arXiv:2606.07590cs.CVcs.AI2026-06

用生物特征分数指导病理模型预训练数据筛选,提升可解释性与效率。

SlideCheck: Guiding Self-Supervised Pretraining of Pathology Foundation Models via Dataset Distributions

  • 基于冻结模型特征设计双头MLP,分别评估异常形态与恶性证据。
  • 通过得分注意力一致性挖掘高置信伪标签,构建正样本子集。
  • 可实现接近全量数据的性能,适合需要可审计预训练数据的场景。

病理基础模型在大量组织切片(WSI)衍生的图像块上进行自监督预训练,但数据构建中的标注常为切片级、稀疏或异质。这种不匹配使得难以理解和控制进入预训练的数据所包含的生物学模式。本文提出SlideCheck,一个轻量级的预训练数据引导工具,基于冻结的病理基础模型图像块特征构建。它不作为独立的图像块诊断模型,而是为图像块提供明确的异常与恶性评分,用于组织、过滤和审计病理预训练数据。SlideCheck采用双头MLP分别建模广泛的异常形态与恶性证据;正则化特征空间评分器提供图像块层面证据估计的监督锚点,而得分-注意力一致性结合图像块评分与切片级多实例学习(MIL)注意力,挖掘高置信伪标签。随后利用这些评分构建广义正样本ViT预训练子集:只要异常或恶性证据超过阈值,即选择该图像块。实验表明,SlideCheck定义的数据分布会影响自监督ViT预训练的下游表现,说明生物组成是病理基础模型开发中可调控的重要因素。经筛选的子集可接近全数据性能,表明显式评分的图像块池可能支持更高效、可审计的预训练数据构建。这些发现使SlideCheck成为将大规模、未区分的图像块池转化为可控且可重用的预训练数据集的数据引导与审计层。

原文摘要 · Abstract (English)

Pathology foundation models are pretrained on large streams of WSI-derived patches, while supervision during data construction is often slide-level, sparse, or heterogeneous. This mismatch makes it difficult to understand and control which biological patterns enter the pretraining data. We propose SlideCheck, a lightweight pretraining data guidance tool built on frozen pathology foundation model patch features. Rather than serving as a standalone patch diagnostic model, SlideCheck provides explicit abnormality and malignancy scores for organizing, filtering, and auditing pathology pretraining data. SlideCheck uses a dual-head MLP to separately model broad abnormal morphology and malignant evidence. A regularized feature-space scorer provides a supervised anchor for patch-level evidence estimation, while score-attention agreement combines patch scores with WSI-level MIL attention to mine high-confidence pseudo labels. The same scores are then used to construct broad-positive ViT pretraining subsets, where a patch is selected if either abnormality or malignancy evidence exceeds a threshold. Experiments show that SlideCheck-defined data distributions influence the downstream behavior of self-supervised ViT pretraining, indicating that biological composition is an important controllable factor in pathology foundation model development. Curated subsets can approach full-data performance, suggesting that explicitly scored patch pools may support more efficient and auditable pretraining data construction. These findings position SlideCheck as a data guidance and auditing layer for transforming large, undifferentiated patch pools into controllable and reusable pretraining datasets.

病理分析自监督学习数据筛选可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。