scpFormer用Transformer统一整合单细胞蛋白组数据,打破抗体面板限制。
scpFormer: A Foundation Model for Unified Representation and Integration of the Single-Cell Proteomics

- 用连续序列锚定替代传统索引,实现无离散化的蛋白表达映射。
- 在超3.9亿细胞上预训练,支持跨批次整合与无监督聚类,效果领先。
- 可扩展临床稀疏数据,适用于癌症药物响应预测等多场景应用。
单细胞蛋白组数据整合常受限于靶向抗体面板的碎片化。为解决此问题,我们提出scpFormer,一种基于Transformer的奠基模型,用于单细胞蛋白组学。该模型在超过3.9亿个细胞上进行预训练,将标准的索引式分词替换为连续、序列锚定的方法。通过结合进化尺度建模(ESM)与值感知表达嵌入,动态将可变抗体面板映射至共享语义空间,避免人为离散化。实验表明,scpFormer生成的全局细胞表征在大规模批次整合与无监督聚类任务中表现优异。其开放词汇架构支持体外面板扩展,有助于在稀疏临床数据中重构生物流形。此外,学习到的蛋白共表达逻辑可迁移至批量组学任务,支持如癌症药物反应预测等应用。scpFormer提供了一种灵活、面板无关的框架,助力可扩展生物标志物发现与精准肿瘤学。
原文摘要 · Abstract (English)
The integration of single-cell proteomic data is often hindered by the fragmented nature of targeted antibody panels. To address this limitation, we introduce scpFormer, a transformer-based foundation model designed for single-cell proteomics. Pre-trained on over 390 million cells, scpFormer replaces standard index-based tokenization with a continuous, sequence-anchored approach. By combining Evolutionary Scale Modeling (ESM) with value-aware expression embeddings, it dynamically maps variable panels into a shared semantic space without artificial discretization. We demonstrate that scpFormer generates global cell representations that perform competitively in large-scale batch integration and unsupervised clustering. Moreover, its open-vocabulary architecture facilitates in silico panel expansion, assisting in the reconstruction of biological manifolds in sparse clinical datasets. Finally, this learned protein co-expression logic is transferable to bulk-omics tasks, supporting applications like cancer drug response prediction. scpFormer provides a versatile, panel-agnostic framework to facilitate scalable biomarker discovery and precision oncology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。