基于4.6亿张病理切片训练出开源大模型Phikon-v2,性能媲美私有数据模型。
Phikon-v2, A large and public feature extractor for biomarker prediction

- 用100多个公开队列的4.6亿张切片自监督训练ViT-L模型
- 在8项任务上优于旧版Phikon,达顶尖模型水平,提升1.75AUC
- 适合医疗AI研究者快速构建病理诊断工具
我们整合了超过100个公开队列的组织病理学切片,构建了一个涵盖30多种癌症类型的多样化数据集,包含4.6亿张病理切片。基于此数据集,我们使用DINOv2框架训练了一个大型自监督视觉变换器,并公开发布了一版模型,命名为Phikon-v2。尽管仅在公开病理切片上训练,Phikon-v2仍超越此前发布的Phikon模型,表现与需使用专有数据训练的其他病理学基础模型(FM)相当。我们的基准测试涵盖8项切片级任务,结果在外部验证队列上报告,确保预训练与评估数据无污染。下游训练采用简单但稳健的集成策略,相比单次微调,整体AUC提升1.75(p<0.001)。我们对比了Phikon(ViT-B)和Phikon-v2(ViT-L)与14种不同病理特征提取器,为当前最全面的评估。结果表明,DINOv2在模型与数据联合扩展方面优于iBOT;近期模型缩放努力总体上提升了下游生物标志物预测性能,其中GigaPath和H-Optimus-0(均为含11亿参数的ViT-g)表现突出。然而,最新顶级模型间的统计差异大多不显著,部分在特定任务如MSI预测中甚至不如内部开发的13倍更小模型。尽管最新基础模型在临床部署中存在局限,但其仍为开发更专业化、低成本的病理编码器提供了优质基础,助力人工智能辅助诊断系统发展。
原文摘要 · Abstract (English)
Gathering histopathology slides from over 100 publicly available cohorts, we compile a diverse dataset of 460 million pathology tiles covering more than 30 cancer sites. Using this dataset, we train a large self-supervised vision transformer using DINOv2 and publicly release one iteration of this model for further experimentation, coined Phikon-v2. While trained on publicly available histology slides, Phikon-v2 surpasses our previously released model (Phikon) and performs on par with other histopathology foundation models (FM) trained on proprietary data. Our benchmarks include eight slide-level tasks with results reported on external validation cohorts avoiding any data contamination between pre-training and evaluation datasets. Our downstream training procedure follows a simple yet robust ensembling strategy yielding a +1.75 AUC increase across tasks and models compared to one-shot retraining (p<0.001). We compare Phikon (ViT-B) and Phikon-v2 (ViT-L) against 14 different histology feature extractors, making our evaluation the most comprehensive to date. Our result support evidences that DINOv2 handles joint model and data scaling better than iBOT. Also, we show that recent scaling efforts are overall beneficial to downstream performance in the context of biomarker prediction with GigaPath and H-Optimus-0 (two ViT-g with 1.1B parameters each) standing out. However, the statistical margins between the latest top-performing FMs remain mostly non-significant; some even underperform on specific indications or tasks such as MSI prediction - deposed by a 13x smaller model developed internally. While latest foundation models may exhibit limitations for clinical deployment, they nonetheless offer excellent grounds for the development of more specialized and cost-efficient histology encoders fueling AI-guided diagnostic tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。