用双尺度图文提示提升病理切片分类,减少标注依赖。
ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification
- 设计双尺度图文提示,融合病理先验知识增强视觉语言模型
- 提出原型引导的图像解码器与上下文引导的文本解码器,高效处理海量切片
- 在三个癌症数据集上表现优越,适合少标签病理分析场景
基于多实例学习(MIL)的框架已成为处理数字病理中具有吉字节级分辨率和分层图像上下文的全切片图像(WSI)的主流方法。然而,这些方法严重依赖大量袋级标签,且仅从原始切片中学习,易受数据分布变化影响。近期基于视觉语言模型(VLM)的方法通过在大规模病理图像-文本对上预训练引入语言先验,但此前文本提示未考虑病理先验知识,未能显著提升性能。此外,此类成对数据的收集及预训练过程耗时且资源密集。为解决上述问题,我们提出一种双尺度视觉语言多实例学习(ViLa-MIL)框架用于全切片图像分类。具体而言,我们基于冻结的大语言模型(LLM)设计双尺度视觉描述性文本提示,以有效提升VLM性能;为将VLM高效迁移至处理WSI,图像分支提出原型引导的补丁解码器,通过将相似补丁聚类至同一原型来逐步聚合补丁特征;文本分支引入上下文引导的文本解码器,通过融合多粒度图像上下文增强文本特征。在三个多癌种、多中心的亚型分类数据集上的大量实验验证了ViLa-MIL的优越性。
原文摘要 · Abstract (English)
Multiple instance learning (MIL)-based framework has become the mainstream for processing the whole slide image (WSI) with giga-pixel size and hierarchical image context in digital pathology. However, these methods heavily depend on a substantial number of bag-level labels and solely learn from the original slides, which are easily affected by variations in data distribution. Recently, vision language model (VLM)-based methods introduced the language prior by pre-training on large-scale pathological image-text pairs. However, the previous text prompt lacks the consideration of pathological prior knowledge, therefore does not substantially boost the model's performance. Moreover, the collection of such pairs and the pre-training process are very time-consuming and source-intensive.To solve the above problems, we propose a dual-scale vision-language multiple instance learning (ViLa-MIL) framework for whole slide image classification. Specifically, we propose a dual-scale visual descriptive text prompt based on the frozen large language model (LLM) to boost the performance of VLM effectively. To transfer the VLM to process WSI efficiently, for the image branch, we propose a prototype-guided patch decoder to aggregate the patch features progressively by grouping similar patches into the same prototype; for the text branch, we introduce a context-guided text decoder to enhance the text features by incorporating the multi-granular image contexts. Extensive studies on three multi-cancer and multi-center subtyping datasets demonstrate the superiority of ViLa-MIL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。