arXiv:2507.20056cs.CVcs.AI2025-07

改进Mamba模型,提升医学图像分割的边界精度与细节保留

FaRMamba: Frequency-based learning and Reconstruction aided Mamba for Medical Segmentation

论文配图:FaRMamba: Frequency-based learning and Reconstruction aided Mamba for Medical Segmentation
图 1 · 摘自论文原文
  • 引入多尺度频域重建模块,恢复被削弱的高频特征
  • 通过自监督重建增强编码器,恢复二维空间关联性
  • 在多个医学数据集上超越主流模型,适合高精度分割场景

由于病变边界模糊(LBA)、高频细节丢失(LHD)以及长程解剖结构建模困难(DC-LRSS),医学图像分割仍具挑战。尽管视觉Mamba通过一维因果状态空间有效缓解了长程依赖问题,但其分块令牌化和一维序列化破坏了局部像素邻接关系,并产生低通滤波效应,导致局部高频信息捕捉不足(LHICD)和二维空间结构退化(2D-SSD),进而加剧边界模糊与细节丢失。为此,本文提出FaRMamba,通过两个互补模块显式解决上述问题:多尺度频域变换模块(MSFM)利用小波、余弦和傅里叶变换分离并重建多带谱,恢复被衰减的高频线索;自监督重建辅助编码器(SSRAE)在共享的Mamba编码器上强制进行像素级重建,以恢复完整的二维空间相关性,提升细纹理与全局上下文。在CAMUS超声心动图、基于MRI的小鼠耳蜗及Kvasir-Seg内窥镜数据集上的大量实验表明,FaRMamba持续优于主流的CNN-Transformer混合模型和现有Mamba变体,在不带来显著计算开销的前提下,实现更优的边界准确性、细节保留与全局一致性。本工作为未来分割模型提供了一个灵活的频域感知框架,直接应对医学影像的核心难题。

原文摘要 · Abstract (English)

Accurate medical image segmentation remains challenging due to blurred lesion boundaries (LBA), loss of high-frequency details (LHD), and difficulty in modeling long-range anatomical structures (DC-LRSS). Vision Mamba employs one-dimensional causal state-space recurrence to efficiently model global dependencies, thereby substantially mitigating DC-LRSS. However, its patch tokenization and 1D serialization disrupt local pixel adjacency and impose a low-pass filtering effect, resulting in Local High-frequency Information Capture Deficiency (LHICD) and two-dimensional Spatial Structure Degradation (2D-SSD), which in turn exacerbate LBA and LHD. In this work, we propose FaRMamba, a novel extension that explicitly addresses LHICD and 2D-SSD through two complementary modules. A Multi-Scale Frequency Transform Module (MSFM) restores attenuated high-frequency cues by isolating and reconstructing multi-band spectra via wavelet, cosine, and Fourier transforms. A Self-Supervised Reconstruction Auxiliary Encoder (SSRAE) enforces pixel-level reconstruction on the shared Mamba encoder to recover full 2D spatial correlations, enhancing both fine textures and global context. Extensive evaluations on CAMUS echocardiography, MRI-based Mouse-cochlea, and Kvasir-Seg endoscopy demonstrate that FaRMamba consistently outperforms competitive CNN-Transformer hybrids and existing Mamba variants, delivering superior boundary accuracy, detail preservation, and global coherence without prohibitive computational overhead. This work provides a flexible frequency-aware framework for future segmentation models that directly mitigates core challenges in medical imaging.

医学分割Mamba频域重建边界优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。