构建多分辨率病理图像基础模型,提升跨尺度诊断能力
From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology

- 采用分层注意力机制融合细胞到组织的多尺度信息
- 在6.24亿图块上自监督预训练,实现跨分辨率特征泛化
- 适用于癌症分型、组织表型分析等病理视觉任务
视觉变换器(ViTs)及其分层变体在计算病理学中表现优异,但多数在单分辨率全幻灯片图像(WSI)上预训练,限制了其在任意分辨率下的泛化能力。大像素级WSI本身包含细胞形态、组织结构和全局上下文等多尺度诊断模式,与病理学家阅片方式一致。本文提出多分辨率金字塔变压器(MRPT),通过层次化聚合从细胞到组织再到全幻灯片级别的多分辨率信息。MRPT采用生物合理的连续跨分辨率注意力(CCRA)机制,捕捉尺度无关的交互,并通过跨分辨率嵌入对齐实现多分辨率语义一致性,生成鲁棒且通用的WSI表示。在6.24亿图块、240万区域和3.6万张全幻灯片上以多分辨率自监督方式预训练,学习丰富的粗到细的组织病理学特征。在34个多样化数据集上的大量实验表明,MRPT在癌症亚型分类、组织表型分析和基于幻灯片的视觉问答(VQA)任务中均超越近期基础模型及多模态大语言模型。
原文摘要 · Abstract (English)
Vision Transformers (ViTs) and their hierarchical variants have achieved strong performance in Computational Pathology (CPath). However, most are pre-trained on single-resolution Whole Slide Images (WSIs), limiting their generalization across arbitrary resolutions. Gigapixel WSIs inherently contain diagnostic patterns at multiple scales, including cellular morphologies, tissue architectures, and global context, mirroring how expert pathologists examine WSIs. We introduce Multi-Resolution Pyramid Transformer (MRPT), a model that hierarchically aggregates multi-resolution information from cellular to tissue and WSI levels. MRPT employs a biologically meaningful Consecutive Cross-Resolution Attention (CCRA) mechanism to capture scale-independent interactions and enforces multi-resolution semantic consistency by aligning embeddings across resolutions, yielding robust and generalizable WSI representations. Pre-trained in a multi-resolution self-supervised manner on 624M patches, 2.4M regions, and 36K WSIs, MRPT learns rich coarse-to-fine histopathology features. Extensive experiments on 34 diverse datasets show that MRPT surpasses recent foundation models and Multimodal Large Language Models (MLLMs) in cancer subtype classification, tissue phenotyping, and Visual Question Answering (VQA) for WSI understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。