arXiv:2503.18709cs.CV2025-03被引 4

自动筛选病理切片细粒度数据,提升视觉大模型性能

Revisiting Automatic Data Curation for Vision Foundation Models in Digital Pathology

  • 用层级聚类从3.5亿张切片中自动挑选平衡样本
  • 发现数据量与多样性存在权衡,影响模型表征质量
  • 新采样策略让下游病理任务表现更优,适合医学影像研究者

视觉基础模型(FMs)正加速数字病理算法的发展,推动生物医学研究。这些模型通过自监督方式学习真实患者样本全幻灯片图像(WSIs)中提取的高异质性切片块的组织学特征。其性能受预训练数据规模、多样性和均衡性显著影响。然而,现有数据筛选主要依赖专家在全幻灯片层面的知识,关注疾病分类和组织类型,却忽视了切片级别的细节。本文针对3.5亿张切片块,探索无监督的切片级自动数据清洗方法。具体地,对预提取的切片嵌入应用层次聚类树,实现基于预训练模型嵌入空间的均匀平衡采样。我们进一步发现,数据集在规模与平衡性之间存在权衡,可能损害模型表征质量,并提出针对性的批采样策略缓解此问题。实验表明,该方法在多种临床相关下游任务中显著提升性能。

原文摘要 · Abstract (English)

Vision foundation models (FMs) are accelerating the development of digital pathology algorithms and transforming biomedical research. These models learn, in a self-supervised manner, to represent histological features in highly heterogeneous tiles extracted from whole-slide images (WSIs) of real-world patient samples. The performance of these FMs is significantly influenced by the size, diversity, and balance of the pre-training data. However, data selection has been primarily guided by expert knowledge at the WSI level, focusing on factors such as disease classification and tissue types, while largely overlooking the granular details available at the tile level. In this paper, we investigate the potential of unsupervised automatic data curation at the tile-level, taking into account 350 million tiles. Specifically, we apply hierarchical clustering trees to pre-extracted tile embeddings, allowing us to sample balanced datasets uniformly across the embedding space of the pretrained FM. We further identify these datasets are subject to a trade-off between size and balance, potentially compromising the quality of representations learned by FMs, and propose tailored batch sampling strategies to mitigate this effect. We demonstrate the effectiveness of our method through improved performance on a diverse range of clinically relevant downstream tasks.

数字病理基础模型数据清洗自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。