arXiv:2603.06378cs.CV2026-03

用结构化状态空间模型提升病理切片分析精度

MoEMambaMIL: Structure-Aware Selective State Space Modeling for Whole-Slide Image Analysis

  • 按多分辨率区域组织切片块,保留空间层级关系
  • 在9个任务中表现最优,准确率显著超越现有方法
  • 适合需要精细组织结构分析的医学图像研究者

全幻灯片图像(WSI)分析因图像规模达百万像素级且具有固有的分层多分辨率结构而面临挑战。现有多种实例学习(MIL)方法常将WSI视为无序的切片块集合,难以捕捉全局组织结构与局部细胞模式之间的结构依赖关系。尽管近期的状态空间模型(SSMs)可高效建模长序列,但如何对WSI标记进行结构化以充分挖掘其空间层次仍是一个开放问题。我们提出MoEMambaMIL,一种面向WSI分析的结构感知状态空间模型框架,结合区域嵌套的选通扫描与专家混合(MoE)建模。通过多分辨率预处理,MoEMambaMIL将切片块标记组织为具备区域感知的序列,保持跨分辨率的空间包含关系。在此结构化序列基础上,通过静态、分辨率特异的专家与动态稀疏专家(带学习路由)解耦分辨率感知编码与区域自适应上下文建模。该设计实现了高效长序列建模,并促进在异质诊断模式下的专家专业化。实验表明,MoEMambaMIL在9个下游任务中均取得最佳性能。

原文摘要 · Abstract (English)

Whole-slide image (WSI) analysis is challenging due to the gigapixel scale of slides and their inherent hierarchical multi-resolution structure. Existing multiple instance learning (MIL) approaches often model WSIs as unordered collections of patches, which limits their ability to capture structured dependencies between global tissue organization and local cellular patterns. Although recent State Space Models (SSMs) enable efficient modeling of long sequences, how to structure WSI tokens to fully exploit their spatial hierarchy remains an open problem.We propose MoEMambaMIL, a structure-aware SSM framework for WSI analysis that integrates region-nested selective scanning with mixture-of-experts (MoE) modeling. Leveraging multi-resolution preprocessing, MoEMambaMIL organizes patch tokens into region-aware sequences that preserve spatial containment across resolutions. On top of this structured sequence, we decouple resolution-aware encoding and region-adaptive contextual modeling via a combination of static, resolution-specific experts and dynamic sparse experts with learned routing. This design enables efficient long-sequence modeling while promoting expert specialization across heterogeneous diagnostic patterns. Experiments demonstrate that MoEMambaMIL achieves the best performance across 9 downstream tasks.

病理图像状态空间模型多实例学习医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。