用切片级测试快速筛选病理图像模型,省去大量计算开销。
From Patches to Patients: A study of the tile-to-slide performance transferability in Digital Pathology

- 通过切片级线性探针评估模型性能,替代复杂的全片分析流程。
- 19个模型在42个全片任务和16个切片任务中表现高度相关,验证了有效性。
- 适合临床研究者快速初选模型,尤其在数据量小或资源有限时实用。
基础模型(FMs)近期显著提升了组织病理学中全片图像(WSI)分析的性能,但为特定临床队列选择最优模型仍需多次预处理、耗时的特征提取及多重实例学习(MIL)聚合器训练。本文研究是否可通过高效的切片级线性探针作为全片级性能的可靠代理,从而避免对每个候选编码器运行完整的全片分析流程。我们在42个全片级任务和16个切片级任务上对19个前沿基础模型进行基准测试,使用ABMIL和均值池化聚合方式,比较切片探针指标与全片结果的相关性。结果显示,不同任务难度下切片与全片性能高度相关,表明编码器表示质量是决定全片成功的关键因素。敏感性分析显示,迁移能力在模型间稳定,更受队列规模和每片切片数影响,而非平均任务难度。我们还测量了切片与全片任务中最佳模型的一致性,发现切片基准能可靠筛选出强候选模型。总体而言,切片级基准可作为高效、实用的初始筛选步骤,而全片评估仍是临床任务最终验证的必要环节。
原文摘要 · Abstract (English)
Foundation Models (FMs) have recently redefined the state-of-the-art in histopathology by providing robust representations for whole-slide image (WSI) analysis. However, selecting the optimal foundation model (FM) for a specific clinical cohort currently requires multiple preprocessing steps, followed by computationally expensive feature extraction and the training of a Multiple Instance Learning (MIL) aggregator for every model. In this work, we investigate whether efficient tile-level linear probing can serve as a reliable proxy for slide-level performance, reducing the need to run full slide-level pipelines for every candidate encoder. We benchmark 19 state-of-the-art FMs on 42 slide-level and 16 tile-level tasks, comparing tile probing metrics against slide-level outcomes using ABMIL and Mean Pooling aggregations. We observe a high correlation between tile and slide performance across varying task difficulties, indicating that encoder representation quality is the primary determinant of WSI success. Sensitivity analyses show that transferability is stable across models and is more influenced by cohort sizes and numbers of tiles per slide than by average task difficulty. We also measure the agreement in best performing models between tile and slide-level tasks, showing tile benchmarks reliably shortlist strong candidates. Overall, our study indicates that tile-level benchmarking provides an efficient and practical first step for narrowing down candidate models, while slide-level evaluation remains essential for final validation on clinical tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。