用高频先验和Mamba注意力提升分割边界精度
WaveSeg: Enhancing Segmentation Precision via High-Frequency Prior and Mamba-Driven Spectrum Decomposition
- 在空间与小波域联合优化特征,增强细节保留
- 小波域中通过Mamba机制建模长程依赖,提升高频结构
- 适合需要精细边界分割的医学图像或遥感任务
当前语义分割网络虽依赖强大预训练编码器,但解码器普遍简单,难以平衡语义上下文与细粒度细节。为此,本文提出新解码器WaveSeg,同时在空间域与小波域优化特征。首先从输入图像学习高频分量作为显式先验,在早期强化边界细节;设计多尺度融合机制Dual Domain Operation(DDO),并提出Spectrum Decomposition Attention(SDA)块,利用Mamba的线性复杂度长程建模能力增强高频结构信息;同时采用重参数化卷积保持小波域低频语义完整性。最后通过残差引导融合,在原始分辨率下整合多尺度特征与边界感知表示,生成语义与结构丰富的特征图。大量实验表明,WaveSeg在标准基准上持续优于现有方法,兼具高效性与高精度。
原文摘要 · Abstract (English)
While recent semantic segmentation networks heavily rely on powerful pretrained encoders, most employ simplistic decoders, leading to suboptimal trade-offs between semantic context and fine-grained detail preservation. To address this, we propose a novel decoder architecture, WaveSeg, which jointly optimizes feature refinement in spatial and wavelet domains. Specifically, high-frequency components are first learned from input images as explicit priors to reinforce boundary details at early stages. A multi-scale fusion mechanism, Dual Domain Operation (DDO), is then applied, and the novel Spectrum Decomposition Attention (SDA) block is proposed, which is developed to leverage Mamba's linear-complexity long-range modeling to enhance high-frequency structural details. Meanwhile, reparameterized convolutions are applied to preserve low-frequency semantic integrity in the wavelet domain. Finally, a residual-guided fusion integrates multi-scale features with boundary-aware representations at native resolution, producing semantically and structurally rich feature maps. Extensive experiments on standard benchmarks demonstrate that WaveSeg, leveraging wavelet-domain frequency prior with Mamba-based attention, consistently outperforms state-of-the-art approaches both quantitatively and qualitatively, achieving efficient and precise segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。