融合状态空间与注意力机制,提升乳腺钼靶多视角分类效率
Mammo-Mamba: A Hybrid State-Space and Transformer Architecture with Sequential Mixture of Experts for Multi-View Mammography
- 用序列专家混合机制动态优化特征表示
- 在CBIS-DDSM数据集上各项指标均优于现有模型
- 适合需要高效高精度乳腺癌筛查的医学AI应用
乳腺癌仍是女性癌症死亡的主要原因,尽管计算机辅助诊断(CAD)系统已有进展。多视角乳腺钼靶图像的准确高效解读对早期发现至关重要,推动了人工智能驱动的CAD模型发展。当前主流多视图乳腺钼靶分类模型基于Transformer架构,但其计算复杂度随图像块数量呈平方增长,亟需更高效的替代方案。为此,我们提出Mammo-Mamba,一种将选择性状态空间模型(SSMs)、基于Transformer的注意力机制和专家驱动特征精炼集成的统一框架。Mammo-Mamba通过定制的SecMamba模块引入序列混合专家(SeqMoE)机制,扩展MambaVision骨干网络。SecMamba是改进的MambaVision模块,通过内容自适应特征精炼增强高分辨率乳腺影像的表征学习能力。这些模块嵌入于MambaVision深层阶段,使模型能通过动态专家门控逐步调整特征侧重,有效克服传统Transformer的局限性。在CBIS-DDSM基准数据集上的评估显示,Mammo-Mamba在所有关键指标上均取得更优分类性能,同时保持计算效率。
原文摘要 · Abstract (English)
Breast cancer (BC) remains one of the leading causes of cancer-related mortality among women, despite recent advances in Computer-Aided Diagnosis (CAD) systems. Accurate and efficient interpretation of multi-view mammograms is essential for early detection, driving a surge of interest in Artificial Intelligence (AI)-powered CAD models. While state-of-the-art multi-view mammogram classification models are largely based on Transformer architectures, their computational complexity scales quadratically with the number of image patches, highlighting the need for more efficient alternatives. To address this challenge, we propose Mammo-Mamba, a novel framework that integrates Selective State-Space Models (SSMs), transformer-based attention, and expert-driven feature refinement into a unified architecture. Mammo-Mamba extends the MambaVision backbone by introducing the Sequential Mixture of Experts (SeqMoE) mechanism through its customized SecMamba block. The SecMamba is a modified MambaVision block that enhances representation learning in high-resolution mammographic images by enabling content-adaptive feature refinement. These blocks are integrated into the deeper stages of MambaVision, allowing the model to progressively adjust feature emphasis through dynamic expert gating, effectively mitigating the limitations of traditional Transformer models. Evaluated on the CBIS-DDSM benchmark dataset, Mammo-Mamba achieves superior classification performance across all key metrics while maintaining computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。