评估了10种模型在癌症图像检索中的表现,发现形态学特征仍有根本局限。
Validation of Whole-Slide Foundation Models for Image Retrieval in TCGA Data

- 对比十种方法在TCGA的9387张切片上进行患者级留一评估
- 最佳模型仅达68%准确率,部分亚型为0%
- 强调需多模态融合与器官/诊断定制化策略
基础模型正在重塑计算病理学,但其在全幻灯片图像检索中相对于强基线(基于补丁和监督聚合)的价值尚不明确。我们在涵盖17个器官和60种诊断的TCGA数据集上,对10种管道进行了基准测试,使用患者级留一评估,包含四种预训练的全幻灯片基础模型、基于注意力的监督多实例学习(ABMIL)聚合器以及五种采样密度下的补丁级检索。性能差异在器官和诊断间大于模型架构间。尽管滑片基础模型TITAN整体表现最强,但优势有限;ABMIL与补丁级方法在Top-1和Top-3准确率上相当,无模型始终占优。形态学特征显著的实体接近上限性能,而罕见、异质性及密切相关的亚型仍具挑战性。误分类与已知观察者间差异大的器官一致,表明纯形态学检索存在内在上限。性能主要由补丁级特征表示驱动,滑片级聚合带来的收益有限,说明在许多场景下聚合可能不必要。这些发现反对存在普适最优架构的观点,支持按器官评估、诊断感知或集成策略、更强特征表示及多模态检索框架。值得注意的是,即使最佳模型在TCGA上也仅达到约68% ± 21%的检索准确率,某些亚型所有方法均达0%准确率,凸显形态表征的根本局限,并表明在可临床部署前仍需重大进展。
原文摘要 · Abstract (English)
Foundation models are reshaping computational histopathology, yet their value for whole-slide image retrieval relative to strong patch-based and supervised aggregation baselines remains unclear. We benchmarked ten pipelines on 9,387 diagnostic slides spanning 17 organs and 60 diagnoses from The Cancer Genome Atlas (TCGA) using patient-level leave-one-patient-out evaluation. Methods included four pre-trained slide foundation models, a supervised attention-based multiple instance learning (ABMIL) aggregator on patch embeddings, and patch-level retrieval across five sampling densities. Performance varied more across organs and diagnoses than across architectures. Although the slide foundation model TITAN achieved the strongest overall results, its advantage was modest; ABMIL and patch-based methods reached comparable Top-1 and Top-3 accuracy, with no model consistently dominant. Morphologically distinctive entities approached ceiling performance, while rare, heterogeneous, and closely related subtypes remained challenging. Misclassifications aligned with organs exhibiting known inter-observer variability, suggesting an intrinsic ceiling for morphology-only retrieval. Performance was driven primarily by patch-level feature representations, with limited benefit from slide-level aggregation, indicating aggregation may be unnecessary in many settings. These findings argue against a universally optimal architecture and instead support organ-resolved benchmarking, diagnosis-aware or ensemble strategies, stronger feature representations, and multimodal retrieval frameworks. Notably, even the best model achieved only $\approx 68\% \pm 21\%$ retrieval accuracy on TCGA, and some subtypes showed $0\%$ accuracy across all methods, highlighting fundamental limitations of morphology-based representations and the need for substantial progress before reliable clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。