CytoBERT让单细胞流式数据跨研究分析成为可能。
CytoBERT: A Foundation Model for Cytometry Data

- 用自监督学习在超5000万细胞上预训练,学通不同标记组合间的关联
- 跨异构数据集微调后分类准确率显著提升,验证了迁移可行性
- 开源模型支持可复现分析,适合免疫与临床研究者使用
流式细胞术测量单个细胞的复杂特征(如免疫细胞计数和蛋白表达),广泛应用于免疫研究和临床。但因实验方案和检测标志物差异,数据高度异质且缺乏标准化。尽管机器学习有潜力揭示细胞生物学深层规律,却难以跨研究应用。近年来基础模型进展缓解此问题,但该领域仍缺乏相应方法。为此,我们提出CytoBERT——一个公开可用、开源、开放权重的基础模型,适用于可变标记面板的单细胞流式数据。该模型在大规模流式数据语料库(15个人类数据集,异构标记面板,超过5000万细胞)上通过标记标准化后进行自监督预训练,能够学习细胞内跨标记的可迁移关系。在样本级分类任务中微调后,证明了在异构流式数据集间实现迁移学习是可行的,为可扩展、通用的流式分析提供起点。代码已开源。
原文摘要 · Abstract (English)
Cytometry measures the complex characteristics of single cells (e.g., counts and protein expression of immune cells) and is widely used across immunological research and clinical settings. However, cytometry data is highly heterogeneous and unstandardized due to experimental protocols and the choice of measured features. While machine learning methods hold the potential to gain deeper insights into cell biology, these challenges make them difficult to apply and transfer across studies. Recent advances in foundation models can alleviate these issues, but corresponding approaches are still scarce in this field. To address this, we provide CytoBERT, a publicly available, open-source, open-weight foundation model for single-cell cytometry data with variable marker panels. CytoBERT is pretrained in a self-supervised manner on a large-scale cytometry corpus (15 human datasets with heterogeneous marker panels and more than 50 million cells) curated through marker standardization, enabling it to learn transferable inter-marker relationships within cells. Fine-tuning CytoBERT for sample-level classification demonstrates that transfer learning across heterogeneous cytometry datasets is feasible, providing a starting point for scalable, generalizable cytometry analysis. Code is available at GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。