现有病理模型对扫描仪差异敏感,影响临床可靠性。
Scanner-Induced Domain Shifts Undermine the Robustness of Pathology Foundation Models
- 在5台扫描仪上测试14种模型,分离出设备带来的干扰
- 多数模型嵌入空间含明显扫描仪特征,影响预测校准
- 模型规模或训练数据量不能保证抗设备差异能力
病理基础模型(PFM)在计算病理学中扮演核心角色,旨在为全切片图像(WSI)提供通用特征提取器。尽管基准表现优异,其对真实世界技术域偏移(如扫描仪差异)的鲁棒性仍不明确。我们系统评估了14种PFM在扫描仪引起的变异下的鲁棒性,涵盖先进自监督模型与自然图像预训练基线。基于384张乳腺癌全切片图像(来自5台扫描仪)的多扫描仪数据集,独立分离扫描仪效应。通过无监督嵌入分析与一系列临床病理监督任务评估鲁棒性。结果表明:当前PFM无法抵消扫描仪导致的域偏移,多数模型在嵌入空间中编码显著扫描仪特异性信息。尽管AUC保持稳定,但扫描仪差异系统性改变嵌入空间,影响下游模型预测的校准,产生依赖扫描仪的偏差,可能影响临床应用可靠性。进一步发现,鲁棒性并非仅取决于训练数据规模、模型大小或新旧程度。尽管视觉-语言模型因数据多样性略具优势,但在下游监督任务中表现较差。结论:PFM的开发与评估需超越以准确率为唯一指标,转向对嵌入稳定性与校准性的显式评估与优化。
原文摘要 · Abstract (English)
Pathology foundation models (PFMs) have become central to computational pathology, aiming to offer general encoders for feature extraction from whole-slide images (WSIs). Despite strong benchmark performance, PFM robustness to real-world technical domain shifts, such as variability from whole-slide scanner devices, remains poorly understood. We systematically evaluated the robustness of 14 PFMs to scanner-induced variability, including state-of-the-art models, earlier self-supervised models, and a baseline trained on natural images. Using a multiscanner dataset of 384 breast cancer WSIs scanned on five devices, we isolated scanner effects independently from biological and laboratory confounders. Robustness is assessed via complementary unsupervised embedding analyses and a set of clinicopathological supervised prediction tasks. Our results demonstrate that current PFMs are not invariant to scanner-induced domain shifts. Most models encode pronounced scanner-specific variability in their embedding spaces. While AUC often remains stable, this masks a critical failure mode: scanner variability systematically alters the embedding space and impacts calibration of downstream model predictions, resulting in scanner-dependent bias that can impact reliability in clinical use cases. We further show that robustness is not a simple function of training data scale, model size, or model recency. None of the models provided reliable robustness against scanner-induced variability. While the models trained on the most diverse data, here represented by vision-language models, appear to have an advantage with respect to robustness, they underperformed on downstream supervised tasks. We conclude that development and evaluation of PFMs requires moving beyond accuracy-centric benchmarks toward explicit evaluation and optimisation of embedding stability and calibration under realistic acquisition variability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。