HistoPLUS提升癌组织切片细胞分析,尤其擅长罕见细胞类型识别。
Towards Comprehensive Cellular Characterisation of H&E slides
- 基于10万+核的泛癌数据集训练,覆盖13种细胞类型。
- 检测质量提升5.2%,分类F1分数提高23.7%,参数量仅为五分之一。
- 可识别7种罕见细胞,适用于未训练过的癌症类型迁移研究。
细胞检测、分割与分类是分析血苏木精-伊红(H&E)切片中肿瘤微环境(TME)的关键。现有方法在未充分研究的细胞类型(罕见或未见于公开数据集)上表现不佳,且跨领域泛化能力有限。为解决这些问题,我们提出HistoPLUS,一种先进的细胞分析模型,基于全新构建的包含108,722个核的泛癌数据集,涵盖13种细胞类型进行训练。在4个独立队列的外部验证中,HistoPLUS在检测质量上优于当前最优模型5.2%,整体分类F1得分提升23.7%,同时仅使用其五分之一的参数量。值得注意的是,该模型首次实现了对7种未充分研究细胞类型的分析,并在13种细胞类型中的8种上取得显著改进。此外,我们证明了HistoPLUS能稳健迁移至训练时未见过的两种肿瘤学适应症。为推动更广泛的TME生物标志物研究,我们已将模型权重与推理代码发布于https://github.com/owkin/histoplus/。
原文摘要 · Abstract (English)
Cell detection, segmentation and classification are essential for analyzing tumor microenvironments (TME) on hematoxylin and eosin (H&E) slides. Existing methods suffer from poor performance on understudied cell types (rare or not present in public datasets) and limited cross-domain generalization. To address these shortcomings, we introduce HistoPLUS, a state-of-the-art model for cell analysis, trained on a novel curated pan-cancer dataset of 108,722 nuclei covering 13 cell types. In external validation across 4 independent cohorts, HistoPLUS outperforms current state-of-the-art models in detection quality by 5.2% and overall F1 classification score by 23.7%, while using 5x fewer parameters. Notably, HistoPLUS unlocks the study of 7 understudied cell types and brings significant improvements on 8 of 13 cell types. Moreover, we show that HistoPLUS robustly transfers to two oncology indications unseen during training. To support broader TME biomarker research, we release the model weights and inference code at https://github.com/owkin/histoplus/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。