arXiv:2502.20850cs.CV2025-02被引 2

用视觉语言模型生成可解释的病理切片表示,提升分析效果与可读性。

VLEER: Vision and Language Embeddings for Explainable Whole Slide Image Representation

  • 基于预训练视觉语言模型提取切片级特征,融合图文信息。
  • 在三个病理切片数据集上表现优于传统视觉特征。
  • 通过文本注释实现结果可解释,适合临床辅助诊断场景。

近年来,视觉语言模型(VLMs)在跨模态理解方面展现巨大潜力。在计算病理学中,针对领域特化的VLMs在大量组织病理图像-文本数据上预训练后,在多种下游任务中表现优异。然而,现有研究主要聚焦于预训练过程和片段级应用,未充分挖掘其在全切片图像(WSI)层面的潜力。本研究假设:预训练的VLMs可通过量化特征提取,自然捕捉到具有信息量且可解释的WSI表征。为此,我们提出视觉与语言嵌入用于可解释的全切片图像表示(VLEER),一种旨在利用VLMs进行WSI表征的新方法。我们在三个病理学全切片图像数据集上系统评估了VLEER,证明其在WSI分析任务中优于传统视觉特征。更重要的是,VLEER具备独特可解释性,能借助文本模态提供详细的病理注释,为切片级病理任务输出清晰推理依据。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) have shown remarkable potential in bridging visual and textual modalities. In computational pathology, domain-specific VLMs, which are pre-trained on extensive histopathology image-text datasets, have succeeded in various downstream tasks. However, existing research has primarily focused on the pre-training process and direct applications of VLMs on the patch level, leaving their great potential for whole slide image (WSI) applications unexplored. In this study, we hypothesize that pre-trained VLMs inherently capture informative and interpretable WSI representations through quantitative feature extraction. To validate this hypothesis, we introduce Vision and Language Embeddings for Explainable WSI Representation (VLEER), a novel method designed to leverage VLMs for WSI representation. We systematically evaluate VLEER on three pathological WSI datasets, proving its better performance in WSI analysis compared to conventional vision features. More importantly, VLEER offers the unique advantage of interpretability, enabling direct human-readable insights into the results by leveraging the textual modality for detailed pathology annotations, providing clear reasoning for WSI-level pathology downstream tasks.

病理分析可解释AI视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。