用线性时间模型实现多尺度病理切片分析,效率更高且精度提升
A Multi-scale Linear-time Encoder for Whole-Slide Image Analysis
- 基于纯Mamba架构并行处理多分辨率切片,构建粗到细推理机制
- 在5个数据集上最高提升6.9% AUC、20.3%准确率和2.3% C-index
- 适合需要高效处理高分辨率病理图像的研究者与临床应用
我们提出Multi-scale Adaptive Recurrent Biomedical Linear-time Encoder(MARBLE),首个纯基于Mamba的多状态多实例学习(MIL)框架,用于全切片图像(WSI)分析。MARBLE并行处理多个放大倍数,并在线性时间状态空间模型中整合从粗到细的推理能力,以极少参数开销高效捕捉跨尺度依赖关系。由于全切片图像具有吉比特级分辨率和分层放大特性,分析难度大;而现有MIL方法通常仅在单尺度运行,基于Transformer的方法则面临二次复杂度的注意力开销。通过结合并行多尺度处理与线性时间序列建模,MARBLE为注意力架构提供了可扩展、模块化的替代方案。在五个公开数据集上的实验表明,其性能最高提升6.9% AUC、20.3%准确率和2.3% C-index,确立了其在多尺度WSI分析中的高效性与泛化能力。
原文摘要 · Abstract (English)
We introduce Multi-scale Adaptive Recurrent Biomedical Linear-time Encoder (MARBLE), the first \textit{purely Mamba-based} multi-state multiple instance learning (MIL) framework for whole-slide image (WSI) analysis. MARBLE processes multiple magnification levels in parallel and integrates coarse-to-fine reasoning within a linear-time state-space model, efficiently capturing cross-scale dependencies with minimal parameter overhead. WSI analysis remains challenging due to gigapixel resolutions and hierarchical magnifications, while existing MIL methods typically operate at a single scale and transformer-based approaches suffer from quadratic attention costs. By coupling parallel multi-scale processing with linear-time sequence modeling, MARBLE provides a scalable and modular alternative to attention-based architectures. Experiments on five public datasets show improvements of up to \textbf{6.9\%} in AUC, \textbf{20.3\%} in accuracy, and \textbf{2.3\%} in C-index, establishing MARBLE as an efficient and generalizable framework for multi-scale WSI analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。