通过频域建模提升医学影像分割精度,兼顾全局与细节。
FEFormer: Frequency-enhanced Vision Transformer for Generic Knowledge Extraction and Adaptive Feature Fusion in Volumetric Medical Image Segmentation

- 引入频域注意力与分解式MLP,同时捕捉局部细节与全局上下文
- 在四个数据集上超越现有方法,计算效率更高
- 适合需要精细结构分割的临床医学影像分析任务
准确分割医学图像中的器官和病灶对临床诊断、预后评估和治疗规划至关重要。尽管视觉变压器(ViTs)表现出色,但在模块与架构设计上仍存在挑战:自注意力难以捕捉细微局部特征,标准MLP缺乏显式空间信息保持机制,传统编码器-解码器依赖简单特征融合策略,无法处理大语义差异,且缺乏将低层信息从编码器传递到解码器的明确机制。为此,我们提出频率增强型视觉变压器(FEFormer),通过显式建模频率信息,联合捕获全局上下文与精细结构细节。FEFormer包含四个新组件:频域增强动态自注意力(FDSA)模块,通过保持局部性的卷积与频域注意力联合捕捉细粒度局部细节与长程依赖;频域分解门控MLP(FGMLP),自适应建模低频与高频成分以增强语义与结构表征;小波引导自适应特征融合(WAFF)模块,在频域实现语义一致的编码器-解码器特征融合;频率使能跨尺度茎桥(FCSB),增强跨尺度的低层特征传播。在四个不同的体数据医学图像分割任务上评估,相比最先进方法,FEFormer实现了更优的分割性能与高计算效率。
原文摘要 · Abstract (English)
Accurate segmentation of organs and lesions in medical images is essential for clinical applications including diagnosis, prognosis, and treatment planning. While Vision Transformers (ViTs) have shown impressive segmentation performance, they face key challenges in module and architecture design. Specifically, self-attention struggles to capture fine-grained local features critical for understanding detailed anatomical structures, standard MLP modules lack explicit mechanisms to preserve spatial information, conventional encoder-decoder architectures rely on naive feature fusion strategies that cannot handle large semantic discrepancies, and existing designs lack explicit mechanisms to propagate low-level information from encoder to decoder. To address these limitations, we propose a Frequency-enhanced Vision Transformer (FEFormer) for robust and efficient volumetric medical image segmentation that explicitly models frequency information to jointly capture global context and fine structural details. FEFormer comprises four novel components: a Frequency-enhanced Dynamic Self-Attention (FDSA) module that jointly captures fine-grained local details and global long-range dependencies through locality-preserving convolution with frequency-domain attention; a Frequency-decomposed Gating MLP (FGMLP) that adaptively models low- and high-frequency components for enhanced semantic and structural representation; a Wavelet-guided Adaptive Feature Fusion (WAFF) module that enables semantically consistent encoder-decoder feature integration in the frequency domain; and a Frequency-enabled Cross-scale Stem Bridge (FCSB) that enhances low-level feature propagation across scales. Evaluated on four diverse volumetric medical image segmentation tasks, FEFormer achieved superior segmentation performance with high computational efficiency compared to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。