融合Transformer与CNN的模型提升结肠息肉分割精度与抗干扰能力
Hybrid(Transformer+CNN)-based Polyp Segmentation
- 结合Transformer长程建模与CNN局部特征提取优势
- 召回率提升1.76%至0.9555,准确率提升0.07%至0.9849
- 适合处理边界模糊及光照、伪影等复杂内窥镜场景
结肠镜检查仍是结肠息肉检测与分割的主要方法。近年来,U-Net、ResUNet、Swin-UNet和PraNet等深度学习网络在息肉分割任务中表现优异。然而,由于息肉尺寸、形状差异大,内窥镜类型、光照条件、成像协议多样,且边界模糊(如黏液、褶皱)等问题,精准分割仍极具挑战。为应对这些难题,本文提出一种混合(Transformer + CNN)模型,通过边界感知注意力机制提升对边界模糊息肉的分割能力,并增强对常见内窥镜伪影(如反光、运动模糊、液体遮挡)的鲁棒性。定量评估显示,该模型在分割精度上显著优于现有方法:召回率提升1.76%(达0.9555),准确率提升0.07%(达0.9849),对各类伪影也表现出更强的适应性。
原文摘要 · Abstract (English)
Colonoscopy is still the main method of detection and segmentation of colonic polyps, and recent advancements in deep learning networks such as U-Net, ResUNet, Swin-UNet, and PraNet have made outstanding performance in polyp segmentation. Yet, the problem is extremely challenging due to high variation in size, shape, endoscopy types, lighting, imaging protocols, and ill-defined boundaries (fluid, folds) of the polyps, rendering accurate segmentation a challenging and problematic task. To address these critical challenges in polyp segmentation, we introduce a hybrid (Transformer + CNN) model that is crafted to enhance robustness against evolving polyp characteristics. Our hybrid architecture demonstrates superior performance over existing solutions, particularly in addressing two critical challenges: (1) accurate segmentation of polyps with ill-defined margins through boundary-aware attention mechanisms, and (2) robust feature extraction in the presence of common endoscopic artifacts, including specular highlights, motion blur, and fluid occlusions. Quantitative evaluations reveal significant improvements in segmentation accuracy (Recall improved by 1.76%, i.e., 0.9555, accuracy improved by 0.07%, i.e., 0.9849) and artifact resilience compared to state-of-the-art polyp segmentation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。