首个病理基础模型综合评测基准,助力精准肿瘤学落地
PathBench: A comprehensive comparison benchmark for pathology foundation models towards precision oncology
- 构建多中心真实数据集,严格防数据泄露
- 覆盖64项诊断与预后任务,验证19个模型表现
- 推动病理大模型从研究走向临床应用
病理基础模型(PFM)正推动计算组织病理学发展,实现高精度、泛化的全切片图像分析,提升癌症诊断与预后评估能力。然而,其临床转化面临模型性能因癌种而异、评估中存在数据泄露、缺乏标准化评测等挑战。现有基准多局限于单一癌种、存在预训练数据重叠或任务覆盖不全。本文提出PathBench,首个综合性评测基准,涵盖10家医院的15,888张全切片图像(来自8,549名患者),覆盖超过64项诊断与预后任务,数据来自私有医疗机构且严格排除预训练使用,避免泄露风险。通过自动化排行榜系统持续评估模型表现。当前对19个PFM的评估显示,Virchow2和H-Optimus-1整体表现最优。该平台为研究人员提供可靠模型开发工具,也为临床医生提供跨场景性能参考,加速病理大模型在常规病理实践中的应用。
原文摘要 · Abstract (English)
The emergence of pathology foundation models has revolutionized computational histopathology, enabling highly accurate, generalized whole-slide image analysis for improved cancer diagnosis, and prognosis assessment. While these models show remarkable potential across cancer diagnostics and prognostics, their clinical translation faces critical challenges including variability in optimal model across cancer types, potential data leakage in evaluation, and lack of standardized benchmarks. Without rigorous, unbiased evaluation, even the most advanced PFMs risk remaining confined to research settings, delaying their life-saving applications. Existing benchmarking efforts remain limited by narrow cancer-type focus, potential pretraining data overlaps, or incomplete task coverage. We present PathBench, the first comprehensive benchmark addressing these gaps through: multi-center in-hourse datasets spanning common cancers with rigorous leakage prevention, evaluation across the full clinical spectrum from diagnosis to prognosis, and an automated leaderboard system for continuous model assessment. Our framework incorporates large-scale data, enabling objective comparison of PFMs while reflecting real-world clinical complexity. All evaluation data comes from private medical providers, with strict exclusion of any pretraining usage to avoid data leakage risks. We have collected 15,888 WSIs from 8,549 patients across 10 hospitals, encompassing over 64 diagnosis and prognosis tasks. Currently, our evaluation of 19 PFMs shows that Virchow2 and H-Optimus-1 are the most effective models overall. This work provides researchers with a robust platform for model development and offers clinicians actionable insights into PFM performance across diverse clinical scenarios, ultimately accelerating the translation of these transformative technologies into routine pathology practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。