让病理图像与文本精准对齐,实现细粒度智能诊断
PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology
- 将整张病理图描述拆解为区域子描述,生成文本引导的区域嵌入
- 在4个基准上超越现有模型,用更少数据达到更高精度
- 适合临床医生做病变定位、分类和指令式查询
尽管视觉-语言模型在计算病理学领域取得显著进展,但全切片图像(WSI)的吉字节级规模与空间异质性仍给多模态理解带来挑战。现有对齐方法难以捕捉数千张图像块中文本描述与视觉线索之间的细粒度对应关系,影响下游任务表现。本文提出PathFLIP(病理细粒度语言-图像预训练),一种面向完整WSI理解的新框架。PathFLIP将滑片级描述分解为区域级子描述,并生成文本条件化的区域嵌入,以促进精确的视觉-语言对齐。通过利用大语言模型(LLMs),PathFLIP可无缝响应多样临床指令,适应不同诊断场景。此外,其具备跨多种范式的通用能力,高效完成滑片级分类与检索、细粒度病灶定位及指令跟随。大量实验表明,PathFLIP在四个代表性基准上优于现有大规模病理视觉-语言模型,同时显著减少训练数据需求,为临床实践中细粒度、指令感知的WSI解析铺平道路。
原文摘要 · Abstract (English)
While Vision-Language Models (VLMs) have achieved notable progress in computational pathology (CPath), the gigapixel scale and spatial heterogeneity of Whole Slide Images (WSIs) continue to pose challenges for multimodal understanding. Existing alignment methods struggle to capture fine-grained correspondences between textual descriptions and visual cues across thousands of patches from a slide, compromising their performance on downstream tasks. In this paper, we propose PathFLIP (Pathology Fine-grained Language-Image Pretraining), a novel framework for holistic WSI interpretation. PathFLIP decomposes slide-level captions into region-level subcaptions and generates text-conditioned region embeddings to facilitate precise visual-language grounding. By harnessing Large Language Models (LLMs), PathFLIP can seamlessly follow diverse clinical instructions and adapt to varied diagnostic contexts. Furthermore, it exhibits versatile capabilities across multiple paradigms, efficiently handling slide-level classification and retrieval, fine-grained lesion localization, and instruction following. Extensive experiments demonstrate that PathFLIP outperforms existing large-scale pathological VLMs on four representative benchmarks while requiring significantly less training data, paving the way for fine-grained, instruction-aware WSI interpretation in clinical practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。