arXiv:2503.20190cs.CV2025-03被引 1

用图文对比学习,无监督构建病理切片统一表征

Cross-Modal Prototype Allocation: Unsupervised Slide Representation Learning via Patch-Text Contrast in Computational Pathology

  • 引入大模型生成原型描述文本,实现图像块与文本的跨模态对齐
  • 无需参数的注意力聚合策略,使切片表示适用于多种下游任务
  • 在4个公开数据集上表现超越现有无监督方法,媲美弱监督模型

随着病理学基础模型的发展,全切片图像(WSI)的表征学习日益受到关注。现有研究虽开发了高质量的图像块特征提取器并设计了精心的聚合方案以生成切片级表征,但主流弱监督方法多基于多重实例学习(MIL),针对特定下游任务定制,通用性受限。部分研究探索无监督学习,但仅关注图像块的视觉模态,忽略了文本数据中蕴含的丰富语义信息。本文提出ProAlign框架,一种跨模态无监督切片表征学习方法。首先利用大语言模型(LLM)为切片中的原型类型生成描述性文本,通过图像块-文本对比构建初始原型嵌入;进一步提出无需参数的注意力聚合策略,基于图像块与原型间的相似性生成无监督切片嵌入,可广泛应用于各类下游任务。在四个公开数据集上的大量实验表明,ProAlign优于现有无监督框架,性能接近部分弱监督模型。

原文摘要 · Abstract (English)

With the rapid advancement of pathology foundation models (FMs), the representation learning of whole slide images (WSIs) attracts increasing attention. Existing studies develop high-quality patch feature extractors and employ carefully designed aggregation schemes to derive slide-level representations. However, mainstream weakly supervised slide representation learning methods, primarily based on multiple instance learning (MIL), are tailored to specific downstream tasks, which limits their generalizability. To address this issue, some studies explore unsupervised slide representation learning. However, these approaches focus solely on the visual modality of patches, neglecting the rich semantic information embedded in textual data. In this work, we propose ProAlign, a cross-modal unsupervised slide representation learning framework. Specifically, we leverage a large language model (LLM) to generate descriptive text for the prototype types present in a WSI, introducing patch-text contrast to construct initial prototype embeddings. Furthermore, we propose a parameter-free attention aggregation strategy that utilizes the similarity between patches and these prototypes to form unsupervised slide embeddings applicable to a wide range of downstream tasks. Extensive experiments on four public datasets show that ProAlign outperforms existing unsupervised frameworks and achieves performance comparable to some weakly supervised models.

病理图像无监督学习跨模态原型学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。