arXiv:2504.05186cs.CVcs.LG2025-04被引 27

用极少病理切片训练出顶尖模型,性能媲美大模型

Training state-of-the-art pathology foundation models with orders of magnitude less data

  • 改进DINOv2框架,适配病理图像训练
  • 仅用1.2万张切片(少两个数量级)即达顶尖性能
  • 适合数据稀缺但追求高性能的病理AI研究者

计算病理学领域近年来得益于现代视觉基础模型(FMs)的发展而迅速进步,这些模型通常在海量病理图像上训练。研究表明,扩大训练数据集、增大模型规模并结合领域特定图像处理技术可显著提升下游任务表现。本文基于此,对标准DINOv2框架进行多项近期改进,以优化病理基础模型的训练。同时引入后训练流程,在高分辨率图像上微调模型,进一步丰富嵌入表征的信息。我们提出了三个新型病理基础模型,其训练所用全视野切片(WSIs)数量比现有最先进模型少两个数量级,但在下游任务中仍表现出相当或更优的性能。即使仅使用TCGA数据(12,000张WSIs)训练的模型,也超越了多数现有模型,平均性能与目前第二好的Virchow2模型相当。这表明,当前病理基础模型的算法仍有巨大优化空间,可充分挖掘大规模数据集的潜力。

原文摘要 · Abstract (English)

The field of computational pathology has recently seen rapid advances driven by the development of modern vision foundation models (FMs), typically trained on vast collections of pathology images. Recent studies demonstrate that increasing the training data set and model size and integrating domain-specific image processing techniques can significantly enhance the model's performance on downstream tasks. Building on these insights, our work incorporates several recent modifications to the standard DINOv2 framework from the literature to optimize the training of pathology FMs. We also apply a post-training procedure for fine-tuning models on higher-resolution images to further enrich the information encoded in the embeddings. We present three novel pathology FMs trained on up to two orders of magnitude fewer WSIs than those used to train other state-of-the-art FMs while demonstrating a comparable or superior performance on downstream tasks. Even the model trained on TCGA alone (12k WSIs) outperforms most existing FMs and, on average, matches Virchow2, the second-best FM published to date. This suggests that there still remains a significant potential for further improving the models and algorithms used to train pathology FMs to take full advantage of the vast data collections.

病理分析小样本学习基础模型图像识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。