首个外部队列验证的乳腺癌生存预测模型基准,发现新模型性能提升有限但小模型更高效。
Benchmarking Pathology Foundation Models for Breast Cancer Survival Prediction

- 统一使用切片特征提取与生存建模框架,跨队列评估多个病理基础模型。
- H-optimus-1表现最佳,二代模型普遍优于一代,但性能提升已趋平缓。
- 小型化模型H0-mini参数不足8%,速度更快,效果反而略超大模型。
病理基础模型(PFMs)作为计算病理学的强大预训练编码器,已广泛应用于下游任务的迁移学习。然而,针对临床有意义的生存预测问题,特别是外部验证条件下的系统性比较仍显不足。本研究对广泛使用及近期提出的多种PFMs在乳腺癌生存预测任务中进行了基准测试,基于标准化的切片级特征提取流程与统一的生存建模框架,评估了三个独立临床队列中超过5,400名患者的长期随访数据。模型在一所队列上训练,在两所独立外部队列上验证,严格检验跨数据集泛化能力。结果表明,H-optimus-1性能最强;整体上模型家族呈现持续代际提升,二代模型普遍优于一代。然而,多数近期模型间绝对性能差异较小,提示单纯扩大预训练数据或模型规模带来的收益正在递减。值得注意的是,紧凑型蒸馏模型H0-mini仅使用不到8%的参数,却在特征提取速度上显著更快,并表现出略优于其大型教师模型H-optimus-0的性能。这些结果提供了首个大规模、经外部验证的乳腺癌生存预测模型基准,为临床工作流中高效部署PFMs提供实用指导。
原文摘要 · Abstract (English)
Pathology foundation models (PFMs) have recently emerged as powerful pretrained encoders for computational pathology, enabling transfer learning across a wide range of downstream tasks. However, systematic comparisons of these models for clinically meaningful prediction problems remain limited, especially in the context of survival prediction under external validation. In this study, we benchmark widely used and recently proposed PFMs for breast cancer survival prediction from whole-slide histopathology images. Using a standardized pipeline based on patch-level feature extraction and a unified survival modeling framework, we evaluate model representations across three independent clinical cohorts comprising more than 5,400 patients with long-term follow-up. Models are trained on one cohort and evaluated on two independent external cohorts, enabling a rigorous assessment of cross-dataset generalization. Overall, H-optimus-1 achieves the strongest survival prediction performance. More broadly, we observe consistent generational improvements across model families, with second-generation PFMs outperforming their first-generation counterparts. However, absolute performance differences between many recent PFMs remain modest, suggesting diminishing returns from further scaling of pretraining data or model size alone. Notably, the compact distilled model H0-mini slightly outperforms its larger teacher model H-optimus-0, despite using fewer than 8% of the parameters and enabling significantly faster feature extraction. Together, these results provide the first large-scale, externally validated benchmark of PFMs for breast cancer survival prediction, and offer practical guidance for efficient deployment of PFMs in clinical workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。