用分子信息监督形态,实现海量细胞的高效精准分类。
CytoFormer: A Molecularly Supervised Cell Foundation Model for Histopathology Cell Classification

- 以空间转录组为标签,训练模型关联细胞形态与分子身份。
- 在16个器官上达0.85准确率、0.78宏F1,复现完整组织结构。
- 标签效率高,适合少样本场景下的病理分析与交互学习。
直接从常规苏木精-伊红(H&E)染色中识别细胞类型可实现大规模单细胞分析,但传统依赖病理医生手动标注,耗时费力且对多数细胞类型不可靠。本文改用分子信息监督形态:基于成像空间转录组技术,在同一组织切片上同时获取细胞的分子身份与形态图像,实现配对标注。研究整合了81个来自16个器官的配对切片,通过聚类、标记基因注释、器官级人工审核与质量控制,构建了包含1540万张细胞图像与对应23种细胞类型的大型数据集。在此基础上训练了CytoFormer——一种具有多任务、器官特异性分类头的细胞基础模型。在空间留出的组织上,其整体准确率达0.85,宏F1为0.78,并成功复现了整片组织的架构。其表示能力具有迁移性:冻结编码器后,线性分类头在四个专家标注基准上优于六种病理基础模型,包括未参与预训练的器官与细胞类型。在交互式主动学习中,仅需少量标注即实现正常上皮与肿瘤的区分,F1达0.82,优于最强基线0.13。该模型将配对的H&E与空间转录组转化为可复用、低标签成本的细胞级分析表征。
原文摘要 · Abstract (English)
Identifying cell types directly from routine haematoxylin and eosin (H&E) histology would enable single-cell analysis at scale, but training such models has relied on manual pathologist annotations, which are slow, expensive and unreliable for many cell types. We instead supervise morphology with molecules. Imaging-based spatial transcriptomics profiles individual cells in situ on a section that can afterwards be stained with H&E, so that molecular identity and morphology are observed for the same physical cell. We assembled 81 such paired Xenium sections spanning 16 organs, derived per-cell labels by clustering, marker-gene annotation, organ-wise human review and quality control, and mapped them onto the cell types commonly reported in each organ. This yielded 15.4 million cells, each with a paired H&E image patch and one of 23 cell types, on which we trained CytoFormer, a cell foundation model with a multi-task, per-organ classification head. On spatially held-out tissue CytoFormer reached an accuracy of 0.85 and a macro-F1 of 0.78 across all 16 organs, and its predictions reproduced the tissue architecture of an entire held-out section. The representation also transfers: with the encoder frozen, a linear head on CytoFormer features performed better than six pathology foundation models on four expert-annotated benchmarks, including on organs and cell types that were not part of pretraining. Finally, in an interactive active-learning setting, CytoFormer's embeddings are markedly more label-efficient than existing pathology foundation models, detecting normal epithelium amid look-alike tumour with an F1 of 0.82 from only a few annotations and leading the strongest baseline by 0.13 in F1. CytoFormer turns paired H&E and spatial transcriptomics into a reusable, label-efficient representation for cell-level analysis of routine histology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。