基础模型仍含医院特征,可能影响病理诊断准确性。
Do Histopathological Foundation Models Eliminate Batch Effects? A Comparative Study
- 用大规模图像预训练的病理模型仍保留医院特异性特征
- 这些特征导致分类偏差,且在特征空间中占主导地位
- 适合关注医疗模型泛化能力的研究者阅读
深度学习在计算病理学中取得显著进展,如疾病诊断、生物标志物预测和预后判断。然而,标注数据稀缺以及批次效应(如不同医院间系统性技术差异)严重影响模型鲁棒性和泛化能力。近期基于数百万至数十亿张图像预训练的病理基础模型被报道可提升下游任务表现。但其是否能完全消除批次效应尚未系统评估。本研究实证发现,基础模型的特征嵌入仍包含明显医院标识,可能导致预测偏差和误分类。这些标识无法通过染色归一化方法去除,在特征空间距离中占主导,并在多个主成分中均显著存在。研究为医学基础模型评估提供新视角,推动更稳健的预训练策略与下游预测器发展。
原文摘要 · Abstract (English)
Deep learning has led to remarkable advancements in computational histopathology, e.g., in diagnostics, biomarker prediction, and outcome prognosis. Yet, the lack of annotated data and the impact of batch effects, e.g., systematic technical data differences across hospitals, hamper model robustness and generalization. Recent histopathological foundation models -- pretrained on millions to billions of images -- have been reported to improve generalization performances on various downstream tasks. However, it has not been systematically assessed whether they fully eliminate batch effects. In this study, we empirically show that the feature embeddings of the foundation models still contain distinct hospital signatures that can lead to biased predictions and misclassifications. We further find that the signatures are not removed by stain normalization methods, dominate distances in feature space, and are evident across various principal components. Our work provides a novel perspective on the evaluation of medical foundation models, paving the way for more robust pretraining strategies and downstream predictors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。