用DINO注意力键提升肠镜息肉分割泛化能力
Unlocking Generalization in Polyp Segmentation with DINO Self-Attention "keys"
- 利用DINO自注意力模块的键特征,搭配简单卷积解码器
- 在多中心数据上实现顶尖性能,尤其在数据少时表现优异
- 无需专用结构,适合临床部署与跨域应用
自动息肉分割对提升结直肠癌临床识别至关重要。尽管深度学习技术已被广泛研究,但现有方法在数据有限或挑战性场景下普遍面临泛化能力不足的问题,且多数依赖复杂的任务特定架构。为此,我们提出一种框架,利用DINO自注意力模块中“键”特征的内在鲁棒性进行稳健分割。不同于传统从视觉变换器(ViT)深层提取特征的方法,本方法仅通过简单卷积解码器处理DINO自注意力的键特征,显著提升性能与泛化能力。我们在多中心数据集上采用两种严格协议验证:领域泛化(DG)和极端单领域泛化(ESDG)。综合统计分析表明,该方案达到当前最优(SOTA)性能,尤其在数据稀缺和困难场景中优势明显。不使用息肉专用架构的情况下,超越nnU-Net和UM-Net等成熟模型。此外,我们系统性地评估了DINO框架的演进过程,量化了其架构改进对下游息肉分割性能的具体影响。
原文摘要 · Abstract (English)
Automatic polyp segmentation is crucial for improving the clinical identification of colorectal cancer (CRC). While Deep Learning (DL) techniques have been extensively researched for this problem, current methods frequently struggle with generalization, particularly in data-constrained or challenging settings. Moreover, many existing polyp segmentation methods rely on complex, task-specific architectures. To address these limitations, we present a framework that leverages the intrinsic robustness of DINO self-attention "key" features for robust segmentation. Unlike traditional methods that extract tokens from the deepest layers of the Vision Transformer (ViT), our approach leverages the key features of the self-attention module with a simple convolutional decoder to predict polyp masks, resulting in enhanced performance and better generalizability. We validate our approach using a multi-center dataset under two rigorous protocols: Domain Generalization (DG) and Extreme Single Domain Generalization (ESDG). Our results, supported by a comprehensive statistical analysis, demonstrate that this pipeline achieves state-of-the-art (SOTA) performance, significantly enhancing generalization, particularly in data-scarce and challenging scenarios. While avoiding a polyp-specific architecture, we surpass well-established models like nnU-Net and UM-Net. Additionally, we provide a systematic benchmark of the DINO framework's evolution, quantifying the specific impact of architectural advancements on downstream polyp segmentation performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。