arXiv:2603.12547cs.CV2026-03

用新型解码器提升医学影像分割通用性,兼顾精度与效率

Decoding Matters: Efficient Mamba-Based Decoder with Distribution-Aware Deep Supervision for Medical Image Segmentation

  • 解码器主导设计,融合CAG、VSSM与可变形卷积增强上下文建模
  • 多阶段分布感知损失使模型在9个数据集上达到最优性能
  • 适合需要跨模态泛化的医学图像分割研究者使用

深度学习在医学图像分割中已取得显著进展,能精准识别肿瘤和组织,但多数方法任务特异性强,在不同成像模态间泛化能力有限。许多现有方法侧重编码器,依赖大型预训练骨干网络,导致计算开销大。本文提出一种以解码器为中心的通用2D医学图像分割方法Deco-Mamba,采用类似U-Net的结构,结合CNN-Transformer-CNN混合编码器与创新解码器。解码器集成共注意力门(CAG)、视觉状态空间模块(VSSM)及可变形卷积精炼块,强化多尺度上下文表示。同时引入分窗分布感知KL散度损失,实现多阶段深度监督。在多个医学图像分割基准上进行大量实验,结果表明该方法在9个数据集上均达到领先性能,具备优异泛化能力且模型复杂度适中。

原文摘要 · Abstract (English)

Deep learning has achieved remarkable success in medical image segmentation, often reaching expert-level accuracy in delineating tumors and tissues. However, most existing approaches remain task-specific, showing strong performance on individual datasets but limited generalization across diverse imaging modalities. Moreover, many methods focus primarily on the encoder, relying on large pretrained backbones that increase computational complexity. In this paper, we propose a decoder-centric approach for generalized 2D medical image segmentation. The proposed Deco-Mamba follows a U-Net-like structure with a Transformer-CNN-Mamba design. The encoder combines a CNN block and Transformer backbone for efficient feature extraction, while the decoder integrates our novel Co-Attention Gate (CAG), Vision State Space Module (VSSM), and deformable convolutional refinement block to enhance multi-scale contextual representation. Additionally, a windowed distribution-aware KL-divergence loss is introduced for deep supervision across multiple decoding stages. Extensive experiments on diverse medical image segmentation benchmarks yield state-of-the-art performance and strong generalization capability while maintaining moderate model complexity. The source code will be released upon acceptance.

医学图像分割解码器设计状态空间模型多尺度建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。