arXiv:2509.15482cs.CVcs.AI2025-09被引 3

对比6种病理学大模型的表征结构,发现训练方式不决定相似性。

Comparing Computational Pathology Foundation Models using Representational Similarity Analysis

  • 用神经科学中的表示相似性分析法,比较六种病理模型表征空间
  • 视觉-语言模型表征更紧凑,但所有模型都高度依赖切片特征
  • 染色标准化可降低切片依赖性,提升模型鲁棒性,适合医学影像研究者

基础模型在计算病理学(CPath)中日益发展,有望支持多种下游任务。尽管已有研究评估模型任务性能,但对其学习表征结构与差异了解有限。本文系统分析六种CPath基础模型的表征空间,涵盖视觉-语言对比学习(CONCH、PLIP、KEEP)和自蒸馏(UNI (v2)、Virchow (v2)、Prov-GigaPath)方法。基于TCGA的H&E图像块进行表示相似性分析发现:UNI2和Virchow2表征结构最不同,Prov-GigaPath与其他模型平均相似度最高;相同训练范式(仅视觉或视觉-语言)并未保证更高表征相似性。所有模型表征均呈现高切片依赖性、低疾病依赖性。染色标准化使各模型切片依赖性下降5.5%(CONCH)至20.5%(PLIP)。视觉-语言模型具有相对紧凑的内在维度,而仅视觉模型则更分散。这些结果提示应提升对切片特异性特征的鲁棒性,优化模型集成策略,并揭示训练范式如何塑造表征。该框架可扩展至其他医学影像领域,以促进基础模型的有效开发与部署。

原文摘要 · Abstract (English)

Foundation models are increasingly developed in computational pathology (CPath) given their promise in facilitating many downstream tasks. While recent studies have evaluated task performance across models, less is known about the structure and variability of their learned representations. Here, we systematically analyze the representational spaces of six CPath foundation models using techniques popularized in computational neuroscience. The models analyzed span vision-language contrastive learning (CONCH, PLIP, KEEP) and self-distillation (UNI (v2), Virchow (v2), Prov-GigaPath) approaches. Through representational similarity analysis using H&E image patches from TCGA, we find that UNI2 and Virchow2 have the most distinct representational structures, whereas Prov-Gigapath has the highest average similarity across models. Having the same training paradigm (vision-only vs. vision-language) did not guarantee higher representational similarity. The representations of all models showed a high slide-dependence, but relatively low disease-dependence. Stain normalization decreased slide-dependence for all models by a range of 5.5% (CONCH) to 20.5% (PLIP). In terms of intrinsic dimensionality, vision-language models demonstrated relatively compact representations, compared to the more distributed representations of vision-only models. These findings highlight opportunities to improve robustness to slide-specific features, inform model ensembling strategies, and provide insights into how training paradigms shape model representations. Our framework is extendable across medical imaging domains, where probing the internal representations of foundation models can support their effective development and deployment.

计算病理表征分析大模型医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。