用病理知识引导Mamba模型,让切片分析更聚焦关键诊断区域。
KHiM-Mamba: Injecting Pathology Knowledge into Mamba via Hidden-State Modulation for Whole Slide Image Analysis

- 在Mamba隐藏状态演化中注入病理先验知识,动态调节信息积累
- 11个公开数据集上实现4类任务的最优性能
- 结合大模型生成组织级语义描述,适配多任务病理分析
全切片图像分析通常被建模为多实例学习(MIL),其中实例特征经上下文更新并聚合为整体切片表示,这一过程称为切片编码动态。近年来,选择性状态空间模型(SSM)因其长序列建模能力和线性复杂度,成为有前景的MIL架构。然而,现有基于SSM的MIL方法仅依赖视觉特征进行编码。在大规模全切片图像中,诊断关键区域稀疏分布于大量无关背景之中,纯视觉驱动的选择性动态可能导致状态更新与读出错位,使演化中的状态累积无关证据,稀释关键诊断线索。本文提出知识感知的隐藏状态调制架构(KHiM-Mamba),创新性地利用显式病理先验调控Mamba的核心选择性状态空间机制,引导切片编码动态向诊断有意义的信息积累。具体而言,重新设计原始SSM层,在隐藏状态演化过程中执行知识调制操作,从而指导每一步编码中哪些视觉证据被积累和检索。此外,引入局部自适应词汇检索模块,利用大语言模型为每个图像块生成细粒度、组织特异的语义描述,实现跨多种任务的精准调制。在11个公共基准上的4项任务实验表明,KHiM-Mamba持续达到最先进性能。
原文摘要 · Abstract (English)
Whole slide image analysis is commonly formulated as multiple instance learning (MIL), where instance features are contextually updated and aggregated into a slide representation, a process we term slide encoding dynamics. Recently, selective state-space models (SSM) have emerged as promising MIL architectures due to their long-sequence modeling capability and linear complexity. However, existing SSM-based MIL methods rely solely on visual features during MIL. Meanwhile, in large-scale WSIs, where sparse diagnostically decisive regions are surrounded by abundant irrelevant information, such purely vision-driven selective dynamics can misallocate state updates and readouts, causing the evolving SSM state to accumulate task-irrelevant evidence and dilute critical diagnostic cues over long scan trajectories. In this work, we propose the Knowledge-Aware Hidden-State Modulation architecture (KHiM-Mamba), which innovatively regulates Mamba's core selective state-space mechanism with explicit knowledge priors, steering slide encoding dynamics toward diagnostically meaningful evidence accumulation. Specifically, we redesign the original SSM layer to perform knowledge modulation operations during the evolution of hidden states, thereby guiding what visual evidence is accumulated and retrieved from the hidden state at each encoding step. Furthermore, we additionally introduce a local-adaptive vocabulary retrieval module that uses large language models to assign each patch fine-grained, tissue-specific semantic descriptions, enabling precise modulation across diverse tasks. Experiments on 11 public benchmarks across 4 tasks show that KHiM-Mamba consistently achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。