提出三模块融合的解码器,提升医学图像分割的边缘与结构精度。
Decoding with Structured Awareness: Integrating Directional, Frequency-Spatial, and Structural Attention for Medical Image Segmentation
- 通过方向感知注意力增强关键区域响应
- 融合空频域特征,强化边界与纹理保留
- 多尺度结构掩码优化跳连,减少冗余信息
为解决Transformer解码器在捕捉边缘细节、识别局部纹理和建模空间连续性方面的不足,本文提出一种专用于医学图像分割的新解码框架,包含三个核心模块。首先,自适应交叉融合注意力(ACFA)模块结合通道增强与空间注意力,并引入可学习的三维方向引导(平面、水平、垂直),提升对关键区域和结构方向的敏感性。其次,三域特征融合注意力(TFFA)模块融合空间、傅里叶与小波域特征,实现联合空频表示,强化全局依赖与结构建模能力,同时保持边缘、纹理等局部信息,在复杂模糊边界场景中表现优异。最后,结构感知多尺度掩码模块(SMMM)通过多尺度上下文与结构显著性过滤优化编码器-解码器间的跳连,有效降低特征冗余,提升语义交互质量。三模块协同作用,显著改善肿瘤分割与器官边界提取等高精度任务性能,提升分割准确率与模型泛化能力。实验表明该框架为医学图像分割提供了一种高效实用的解决方案。
原文摘要 · Abstract (English)
To address the limitations of Transformer decoders in capturing edge details, recognizing local textures and modeling spatial continuity, this paper proposes a novel decoder framework specifically designed for medical image segmentation, comprising three core modules. First, the Adaptive Cross-Fusion Attention (ACFA) module integrates channel feature enhancement with spatial attention mechanisms and introduces learnable guidance in three directions (planar, horizontal, and vertical) to enhance responsiveness to key regions and structural orientations. Second, the Triple Feature Fusion Attention (TFFA) module fuses features from Spatial, Fourier and Wavelet domains, achieving joint frequency-spatial representation that strengthens global dependency and structural modeling while preserving local information such as edges and textures, making it particularly effective in complex and blurred boundary scenarios. Finally, the Structural-aware Multi-scale Masking Module (SMMM) optimizes the skip connections between encoder and decoder by leveraging multi-scale context and structural saliency filtering, effectively reducing feature redundancy and improving semantic interaction quality. Working synergistically, these modules not only address the shortcomings of traditional decoders but also significantly enhance performance in high-precision tasks such as tumor segmentation and organ boundary extraction, improving both segmentation accuracy and model generalization. Experimental results demonstrate that this framework provides an efficient and practical solution for medical image segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。