arXiv:2602.16696q-bio.GNcs.LG2026-02

不用复杂模型,简单线性方法也能在单细胞数据上达到顶尖性能

Parameter-free representations outperform single-cell foundation models on downstream benchmarks

  • 用标准化+线性方法替代深度学习模型
  • 在多个基准上表现优于或接近最先进模型
  • 尤其在新细胞类型和物种上更胜一筹,适合生物学家使用

单细胞RNA测序(scRNA-seq)数据具有强而可重复的统计结构。这推动了基于Transformer架构的大规模基础模型(如TranscriptFormer)的发展,其通过将基因嵌入潜在空间来学习基因表达的生成模型。这些嵌入已在细胞类型分类、疾病状态预测和跨物种学习等下游任务中实现最先进(SOTA)性能。本文探讨是否可通过无需计算密集型深度学习的表示方法获得类似效果。我们采用简单、可解释的流程,结合精细归一化与线性方法,在多个常用评估单细胞基础模型的基准上取得了SOTA或近SOTA表现,甚至在训练数据中未包含的新细胞类型和生物体的分布外任务上超越了基础模型。结果表明,细胞身份的生物学特征可通过单细胞基因表达数据的简单线性表示捕捉,强调了严格基准测试的必要性。

原文摘要 · Abstract (English)

Single-cell RNA sequencing (scRNA-seq) data exhibit strong and reproducible statistical structure. This has motivated the development of large-scale foundation models, such as TranscriptFormer, that use transformer-based architectures to learn a generative model for gene expression by embedding genes into a latent vector space. These embeddings have been used to obtain state-of-the-art (SOTA) performance on downstream tasks such as cell-type classification, disease-state prediction, and cross-species learning. Here, we ask whether similar performance can be achieved without utilizing computationally intensive deep learning-based representations. Using simple, interpretable pipelines that rely on careful normalization and linear methods, we obtain SOTA or near SOTA performance across multiple benchmarks commonly used to evaluate single-cell foundation models, including outperforming foundation models on out-of-distribution tasks involving novel cell types and organisms absent from the training data. Our findings highlight the need for rigorous benchmarking and suggest that the biology of cell identity can be captured by simple linear representations of single cell gene expression data.

单细胞线性模型基础模型生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。