arXiv:2511.17355cs.CV2025-11被引 1

统一注意力与Mamba架构,提升肿瘤细胞分类与分割性能

UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification

  • 将注意力与Mamba融合为统一结构,无需手动调比例
  • 细胞分类准确率提升至78%(34.9万细胞),分割精度达80%(406个图像块)
  • 适用于多模态医学图像分析,尤其适合肿瘤细胞研究

受Mamba架构在视觉与语言领域成功启发,我们提出统一注意力-Mamba(UAM)主干网络。与以往固定比例融合注意力与Mamba的混合方法不同,UAM在单一连贯架构中灵活结合两者优势,避免人工比例调优,提升编码能力。我们设计两种UAM变体以全面评估该统一结构的优势。基于此主干,进一步构建多模态UAM框架,实现细胞级分类与图像分割联合任务。实验表明,UAM在公开基准上两项任务均达到领先水平,超越主流基于图像的基础模型。细胞分类准确率从74%提升至78%(样本量n=349,882),肿瘤分割精确度从75%提升至80%(图像块数n=406)。

原文摘要 · Abstract (English)

Inspired by the recent success of the Mamba architecture in vision and language domains, we introduce a Unified Attention-Mamba (UAM) backbone. Unlike previous hybrid approaches that integrate Attention and Mamba modules in fixed proportions, our unified design flexibly combines their capabilities within a single cohesive architecture, eliminating the need for manual ratio tuning and improving encode capability. We develop two UAM variants to comprehensively evaluate the benefits of this unified structure. Building on this backbone, we further propose a multimodal UAM framework that jointly performs cell-level classification and image segmentation. Experimental results demonstrate that UAM achieves state-of-the-art performance across both tasks on public benchmarks, surpassing leading image-based foundation models. It improves cell classification accuracy from 74\% to 78\% ($n$=349,882 cells), and tumor segmentation precision from 75\% to 80\% ($n$=406 patches).

多模态肿瘤分割Mamba细胞分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。