arXiv:2608.05928cs.LG2026-08

用基因块联合嵌入提升单细胞数据表征能力

BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells

论文配图:BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells
图 1 · 摘自论文原文
  • 以蛋白互作和共表达定义基因块,预测其整体表示
  • 在多个任务中表现优于基线模型,误差最低
  • 适合单细胞生物信息学研究者使用

单细胞转录组是协调生物程序的稀疏观测,但大多数自监督模型通过重建单个基因进行学习。本文提出BioM-JEPA,一种联合嵌入预测架构,通过蛋白互作和语料库衍生共表达证据定义图连接的基因块,预测其聚合表示。学生网络从细胞中其余基因推断目标块表示,教师网络则提供完整基因集对应的参考。在测试诊断中,块级预测生成的嵌入有效秩更高,与检测基因深度关联更弱,优于词预测、随机块和重建控制。在CellBench任务中,冻结的BioM-JEPA嵌入保留了表达、通路和邻域信息,扰动响应误差最低。表示诊断结果与胰腺经典程序及遗传扰动组合关系一致。线性注意力避免构建二次基因-基因注意力矩阵;在批次大小为8的一轮匹配实验中,BioM-JEPA的微调吞吐量比scFoundation高5.75倍,保留嵌入吞吐量高3.76倍。这些结果支持图连接基因块作为单细胞生物学中JEPA式表征学习的有效预测单元。

原文摘要 · Abstract (English)

Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes. Here we present BioM-JEPA, a joint-embedding predictive architecture that instead predicts aggregate representations of graph-connected gene blocks defined by protein-association and corpus-derived coexpression evidence. A student network infers each target-block representation from the remaining genes in a cell, while a slowly updated teacher supplies the corresponding target from the full observed gene set. Under the reported extraction procedure, block-level prediction produced embeddings with higher effective rank and weaker association with detected-gene depth in the tested diagnostics than token-prediction, random-block and reconstruction controls. Across CellBench tasks, frozen BioM-JEPA embeddings retained expression, pathway and neighbourhood information and achieved the lowest aggregate perturbation-response error among the evaluated models. Representation diagnostics were also consistent with canonical pancreatic programmes and compositional relationships between genetic perturbations. Linear attention avoids constructing a quadratic gene-by-gene attention matrix; in a matched one-epoch hPancreas experiment at batch size 8, BioM-JEPA provided 5.75-fold higher fine-tuning throughput and 3.76-fold higher held-out embedding throughput than scFoundation. Together, these results support graph-connected gene blocks as useful prediction units for JEPA-style representation learning in single-cell biology.

单细胞基因网络自监督学习嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。