融合Transformer与CNN的模型提升皮肤病变分割精度
Deep Skin Lesion Segmentation with Transformer-CNN Fusion: Toward Intelligent Skin Cancer Analysis
- 用Transformer捕捉全局语义,保留CNN的局部纹理特征
- 边界引导注意力与多尺度上采样提升边缘定位能力
- 在复杂病变场景下表现优异,适合医学图像分析
本文提出一种基于改进TransUNet架构的高精度语义分割方法,以应对皮肤病变图像中复杂的结构、模糊的边界和显著的尺度变化挑战。该方法将Transformer模块融入传统编码器-解码器框架,用于建模全局语义信息,同时保留卷积分支以保持局部纹理和边缘特征,增强对细粒度结构的感知能力。设计了边界引导注意力机制和多尺度上采样路径,进一步提升病变边界的定位精度与分割一致性。通过对比实验、超参数敏感性分析、数据增强效果、输入分辨率变化及训练数据划分比例测试验证了方法的有效性。结果表明,所提模型在mIoU、mDice和mAcc指标上均优于现有代表性方法,展现出更强的病变识别准确性和鲁棒性,尤其在复杂场景下具备更优的边界重建与结构恢复能力,满足自动化皮肤病变分割任务的关键需求。
原文摘要 · Abstract (English)
This paper proposes a high-precision semantic segmentation method based on an improved TransUNet architecture to address the challenges of complex lesion structures, blurred boundaries, and significant scale variations in skin lesion images. The method integrates a transformer module into the traditional encoder-decoder framework to model global semantic information, while retaining a convolutional branch to preserve local texture and edge features. This enhances the model's ability to perceive fine-grained structures. A boundary-guided attention mechanism and multi-scale upsampling path are also designed to improve lesion boundary localization and segmentation consistency. To verify the effectiveness of the approach, a series of experiments were conducted, including comparative studies, hyperparameter sensitivity analysis, data augmentation effects, input resolution variation, and training data split ratio tests. Experimental results show that the proposed model outperforms existing representative methods in mIoU, mDice, and mAcc, demonstrating stronger lesion recognition accuracy and robustness. In particular, the model achieves better boundary reconstruction and structural recovery in complex scenarios, making it well-suited for the key demands of automated segmentation tasks in skin lesion analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。