多分辨率病理-语言模型提升癌症分类准确率
Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
- 在多分辨率全切片图像上对齐图文,构建跨尺度视觉表征
- 在TCGA数据集上用3400万图文对预训练,任务表现超越现有模型
- 适合需要细粒度组织分析的病理诊断与生存预测研究
在计算病理学中,视觉语言模型主要聚焦于单分辨率图像的图文对齐,但难以满足癌症亚型分类、组织表型分析和生存预测等任务对细节的需求。为此,我们提出一种基于全切片图像(WSIs)的多分辨率范式,从不同放大倍数提取组织切片并生成对应文本描述。通过引入多分辨率图文对齐及跨分辨率对齐机制,结合多模态编码器增强上下文捕捉能力,使模型能更全面地提取特征。我们设计了新颖的损失函数以丰富表示能力,提升判别力与跨分辨率泛化性能。模型在包含3400万图像-文本对的TCGA数据集上预训练,微调后在多个数据集和任务中均优于当前最优方法,验证了其有效性。代码已开源。
原文摘要 · Abstract (English)
In Computational Pathology (CPath), the introduction of Vision-Language Models (VLMs) has opened new avenues for research, focusing primarily on aligning image-text pairs at a single magnification level. However, this approach might not be sufficient for tasks like cancer subtype classification, tissue phenotyping, and survival analysis due to the limited level of detail that a single-resolution image can provide. Addressing this, we propose a novel multi-resolution paradigm leveraging Whole Slide Images (WSIs) to extract histology patches at multiple resolutions and generate corresponding textual descriptions through advanced CPath VLM. We introduce visual-textual alignment at multiple resolutions as well as cross-resolution alignment to establish more effective text-guided visual representations. Cross-resolution alignment using a multimodal encoder enhances the model's ability to capture context from multiple resolutions in histology images. Our model aims to capture a broader range of information, supported by novel loss functions, enriches feature representation, improves discriminative ability, and enhances generalization across different resolutions. Pre-trained on a comprehensive TCGA dataset with 34 million image-language pairs at various resolutions, our fine-tuned model outperforms state-of-the-art (SOTA) counterparts across multiple datasets and tasks, demonstrating its effectiveness in CPath. The code is available on GitHub at: https://github.com/BasitAlawode/MR-PLIP
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。