评测病理模型在数据分布变化下的表现,发现图像风格差异比癌症分级分布更影响效果。
Evaluating Computational Pathology Foundation Models for Prostate Cancer Grading under Distribution Shifts
- 用全切片图像对比不同预训练模型在跨机构数据上的表现
- 跨机构迁移时性能显著下降,但标签分布变化影响较小
- 强调下游训练数据质量与多样性对模型泛化至关重要
病理基础模型(PFMs)作为计算病理学的强大预训练编码器已崭露头角,但其在临床相关分布偏移下的鲁棒性仍不明确。我们基于PANDA数据集,在前列腺癌分级任务中评估了近期PFMs的鲁棒性。将PFMs作为冻结的局部切片特征提取器,集成到弱监督的切片级分级模型中,考察两种关键分布偏移的影响:不同采集机构间的全切片图像外观差异,以及癌症分级组标签分布的变化。在同分布设置下,所有PFMs均表现优异,明显优于自然图像预训练基线。然而在从Radboud到Karolinska的跨机构迁移中,所有模型性能均大幅下降,表明大规模预训练本身不足以保证下游泛化能力。相比之下,PFMs对标签分布偏移的敏感性较低,说明视觉层面的域偏移是主要挑战。表示分析进一步证实,各PFMs中机构间始终存在显著的域分离现象。尽管存在与分级相关的结构,但其强度相对较弱,表明学习到的特征空间中域相关变化占主导地位。这些结果为PFMs在分布偏移下的表现提供了全面基准,并揭示了一个重要实践启示:尽管PFMs提供强表征能力,其泛化能力仍受限于下游预测模型训练数据的质量与多样性。
原文摘要 · Abstract (English)
Pathology foundation models (PFMs) have emerged as powerful pretrained encoders for computational pathology, but their robustness under clinically relevant distribution shifts remains insufficiently understood. We benchmark the robustness of recent PFMs in the setting of prostate cancer grading from whole-slide images (WSIs). Using the PANDA dataset, we evaluate PFMs as frozen patch-level feature extractors within weakly supervised slide-level grading models, and assess robustness to two important forms of distribution shift: shifts in WSI image appearance across collection sites, and shifts in the label distribution over cancer grade groups. Across in-distribution settings, PFMs consistently achieve strong performance and clearly outperform a natural-image baseline. Under cross-site transfer from Radboud to Karolinska, however, performance drops substantially for all models, showing that large-scale pretraining alone does not guarantee robust downstream generalization. In contrast, PFMs are less sensitive to label-distribution shift, indicating that visually grounded domain shift is the dominant challenge. Representation analysis further supports these findings by revealing persistent domain separation between sites across all PFMs. While grade-related structure is present, it is comparatively weak, indicating that domain-related variation dominates in the learned feature space. Together, these results provide a comprehensive benchmark of PFMs under distribution shift and highlight an important practical message: although PFMs provide strong representations, generalizability remains constrained by the quality and diversity of the data used to train downstream prediction models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。