用Mamba架构高效分析病理切片,无需预训练即可达到顶尖水平。
From Pixels to Gigapixels: Bridging Local Inductive Bias and Long-Range Dependencies with Pixel-Mamba
- 结合Mamba与局部先验,分层融合局部与全局信息。
- 在肿瘤分期和生存分析任务中超越现有模型,且无需病理预训练。
- 适合需要高效处理超大病理图像的医疗AI研究者使用。
组织病理学在医学诊断中至关重要,全切片图像(WSIs)为临床决策提供关键信息。然而,其巨大尺寸与复杂性对深度学习模型在计算效率和表征学习方面带来挑战。本文提出Pixel-Mamba,一种新型深度学习架构,专为高效处理千兆像素级WSI设计。该模型采用具有线性内存复杂度的状态空间模型(SSM)——Mamba模块,并通过逐步扩展的标记(tokens)引入局部归纳偏置,类似卷积神经网络。这使得Pixel-Mamba能够层次化融合局部与全局信息,同时有效应对计算瓶颈。令人瞩目的是,即使未进行任何病理领域预训练,Pixel-Mamba在多种肿瘤分期与生存分析任务中的表现仍达到或超越当前最先进(SOTA)的基础模型,这些模型通常在数百万张WSI或图文对上预训练。大量实验验证了Pixel-Mamba作为端到端WSI分析强大而高效框架的有效性。
原文摘要 · Abstract (English)
Histopathology plays a critical role in medical diagnostics, with whole slide images (WSIs) offering valuable insights that directly influence clinical decision-making. However, the large size and complexity of WSIs may pose significant challenges for deep learning models, in both computational efficiency and effective representation learning. In this work, we introduce Pixel-Mamba, a novel deep learning architecture designed to efficiently handle gigapixel WSIs. Pixel-Mamba leverages the Mamba module, a state-space model (SSM) with linear memory complexity, and incorporates local inductive biases through progressively expanding tokens, akin to convolutional neural networks. This enables Pixel-Mamba to hierarchically combine both local and global information while efficiently addressing computational challenges. Remarkably, Pixel-Mamba achieves or even surpasses the quantitative performance of state-of-the-art (SOTA) foundation models that were pretrained on millions of WSIs or WSI-text pairs, in a range of tumor staging and survival analysis tasks, {\bf even without requiring any pathology-specific pretraining}. Extensive experiments demonstrate the efficacy of Pixel-Mamba as a powerful and efficient framework for end-to-end WSI analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。