通过语义重排与检索引导状态空间建模,提升病理全切片图像分析精度与效率。
SemaMIL: Semantic-Aware Multiple Instance Learning with Retrieval-Guided State Space Modeling for Whole Slide Images
- 基于语义聚类重排切片块顺序,保留组织学上下文关系。
- 在4个数据集上优于主流方法,参数和计算量更少。
- 适合需要高精度与可解释性的医学图像分析研究者。
多实例学习(MIL)已成为计算病理学中从全切片图像(WSI)提取判别特征的主流方法。基于注意力的MIL方法虽能识别关键区域,但常忽略上下文关系;Transformer模型虽能建模交互,却需二次计算开销且易过拟合;状态空间模型(SSMs)具有线性复杂度,但打乱切片块顺序会破坏组织学意义并降低可解释性。本文提出SemaMIL,融合语义重排(SR),通过可逆排列自适应地将语义相似的切片块排序;引入语义引导检索状态空间模块(SRSM),选取代表性查询来调整状态空间参数,增强全局建模能力。在4个WSI亚型数据集上的评估表明,相比强基线,SemaMIL在更少的浮点运算数(FLOPs)和参数量下达到最优准确率。
原文摘要 · Abstract (English)
Multiple instance learning (MIL) has become the leading approach for extracting discriminative features from whole slide images (WSIs) in computational pathology. Attention-based MIL methods can identify key patches but tend to overlook contextual relationships. Transformer models are able to model interactions but require quadratic computational cost and are prone to overfitting. State space models (SSMs) offer linear complexity, yet shuffling patch order disrupts histological meaning and reduces interpretability. In this work, we introduce SemaMIL, which integrates Semantic Reordering (SR), an adaptive method that clusters and arranges semantically similar patches in sequence through a reversible permutation, with a Semantic-guided Retrieval State Space Module (SRSM) that chooses a representative subset of queries to adjust state space parameters for improved global modeling. Evaluation on four WSI subtype datasets shows that, compared to strong baselines, SemaMIL achieves state-of-the-art accuracy with fewer FLOPs and parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。