构建首个针对细胞表型分析的综合基准,揭示现有模型在真实病理图像上的性能短板。
PhenoBench: A Comprehensive Benchmark for Cell Phenotyping
- 提出PhenoBench基准,包含14类精细标注的H&E染色图像数据集
- 模型在新数据集上宏平均F1低至0.20,远低于旧基准表现
- 适合评估模型在真实病理场景下的泛化能力,尤其关注技术与医学域偏移
数字病理学中涌现出大量基础模型(FM),但其在细胞表型分析上的性能尚未得到统一评估。为此,我们提出PhenoBench:一个基于苏木精-伊红(H&E)染色病理图像的细胞表型综合基准。我们提供了新的PhenoCell数据集,包含14类通过多重成像识别的精细细胞类型,并提供即用型微调与评估代码,支持在不同泛化场景下系统评估多个主流病理基础模型在密集细胞表型预测上的表现。我们对现有模型进行了广泛基准测试,揭示了其在技术与医学领域偏移下的泛化行为。尽管这些模型在Lizard和PanNuke等已有基准上取得宏平均F1 > 0.70的成绩,但在PhenoCell上得分低至0.20,表明任务难度远超以往基准未捕捉的程度,确立了PhenoCell作为未来基础模型与监督模型基准测试的核心资产。代码与数据已在GitHub公开。
原文摘要 · Abstract (English)
Digital pathology has seen the advent of a wealth of foundational models (FM), yet to date their performance on cell phenotyping has not been benchmarked in a unified manner. We therefore propose PhenoBench: A comprehensive benchmark for cell phenotyping on Hematoxylin and Eosin (H&E) stained histopathology images. We provide both PhenoCell, a new H&E dataset featuring 14 granular cell types identified by using multiplexed imaging, and ready-to-use fine-tuning and benchmarking code that allows the systematic evaluation of multiple prominent pathology FMs in terms of dense cell phenotype predictions in different generalization scenarios. We perform extensive benchmarking of existing FMs, providing insights into their generalization behavior under technical vs. medical domain shifts. Furthermore, while FMs achieve macro F1 scores > 0.70 on previously established benchmarks such as Lizard and PanNuke, on PhenoCell, we observe scores as low as 0.20. This indicates a much more challenging task not captured by previous benchmarks, establishing PhenoCell as a prime asset for future benchmarking of FMs and supervised models alike. Code and data are available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。