arXiv:2608.00985cs.LGcs.AI2026-08

通过互补转录组视图对比学习,提升单细胞表示能力。

Beyond Gene Reconstruction: Learning Cell Representations through Complementary Transcriptomic Views

论文配图:Beyond Gene Reconstruction: Learning Cell Representations through Complementary Transcriptomic Views
图 1 · 摘自论文原文
  • 基于共表达结构分割基因,构建细胞的两个互补视图。
  • 在六网络基因调控网评估中,平均AUROC和AUPRC均领先。
  • 适合需要高质量细胞表征的下游任务研究者。

单细胞转录组数据的快速增长推动了以掩码表达值重建为主的基础模型发展。然而,该目标虽促进基因依赖性学习,却未直接优化全细胞表征,而后者对下游任务至关重要。为此,我们提出一种对比预训练框架,通过互补转录组视图学习细胞表征。针对标准对比学习难以直接应用于单细胞场景的问题,我们在三个维度进行适配:共表达引导的基因分块、表达感知的负样本构造、以及能力感知的对比启动机制。首先,根据基因共表达结构对每细胞基因进行划分,生成两个互补视图;其次,为防止模型利用基因集身份作为捷径,通过置换表达值但保留基因身份的方式构造硬负样本;最后,引入能力感知控制器决定对比目标的应用时机。在细胞类型注释与基因调控网络推断任务上的实验表明,该方法在评估协议下具有竞争力。在六网络基因调控网络(GRN)评估中,本方法在平均AUROC和AUPRC点估计上均优于对比模型,且各网络最优表现模型略有差异。结果确立了互补视图对比学习是超越基因重建的单细胞预训练有效方向。

原文摘要 · Abstract (English)

The rapid growth of single-cell transcriptomic data has enabled the development of foundation models pretrained primarily by reconstructing masked expression values. This objective encourages these models to learn gene dependencies but does not directly optimize whole-cell representations, which are essential for many downstream tasks. To bridge this gap, we propose a contrastive pretraining framework that learns cell representations through complementary transcriptomic views. Since standard contrastive learning is not readily applicable to single-cell pretraining, we introduce specific adaptations along three dimensions --- co-expression-guided gene partitioning, expression-aware contrast-set construction, and competence-gated contrastive onset. Specifically, we first construct two complementary views of each cell by partitioning its genes according to their co-expression structure. Then, to prevent the model from using gene-set identity as a shortcut, we construct hard negatives by permuting expression values while keeping gene identities unchanged. Finally, we introduce a competence-aware controller to determine how the contrastive objective is applied. Experiments on cell-type annotation and gene regulatory network inference demonstrate competitive transfer under the evaluated protocols. In the six-network GRN evaluation, our method records the highest mean AUROC and AUPRC point estimates among the compared variants, while the highest-scoring variant differs across individual networks. These results establish complementary-view contrastive learning as an effective direction for single-cell pretraining beyond gene reconstruction.

单细胞对比学习表征学习基因调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。