arXiv:2506.14861q-bio.GNcs.AI2025-06被引 3

用全细胞表达解码提升转录组模型的细胞表征能力

BMFM-RNA: whole-cell expression decoding improves transcriptomic foundation models

  • 通过单个CLS token重建全基因集,构建信息密集瓶颈
  • 在所有下游任务中优于传统掩码语言建模,尽管训练时重构误差更高
  • 支持零样本多粒度细胞注释,适合生物医学基础模型研发

以掩码语言建模预训练的转录组基础模型虽能实现低预训练损失,但下游任务中的细胞表征效果不佳。本文提出全细胞表达解码(WCED),让模型仅从一个CLS token嵌入中重构全部基因表达谱,即使输入有限也能形成信息量最大化瓶颈。尽管训练时重构误差更高,WCED在所有下游指标上均优于MLM。基因级误差分析显示,两种方法都更倾向于学习与稳定转录程序共变的基因,而非受瞬时因素驱动的基因。此外,引入利用细胞本体结构的分层交叉熵损失,实现多粒度零样本注释。结合这些目标训练的模型在CZI基准测试中表现最优,适用于零样本批次整合与线性探测细胞类型注释。相关方法已开源至biomed-multi-omic框架。

原文摘要 · Abstract (English)

Transcriptomic foundation models pretrained with masked language modeling can achieve low pretraining loss yet produce poor cell representations for downstream tasks. We introduce whole-cell expression decoding (WCED), where models reconstruct the entire gene vocabulary from a single CLS token embedding, even with limited inputs, creating a maximally informative bottleneck. WCED consistently outperforms MLM on all downstream metrics despite higher reconstruction error during training. Gene-level error tracking reveals that both methods preferentially learn genes whose expression co-varies with stable transcriptional programs rather than those driven by transient factors. We further add hierarchical cross-entropy loss that exploits Cell Ontology structure for zero-shot annotation at multiple granularity levels. Models trained with these objectives achieve best overall performance across CZI benchmarks, on zero-shot batch integration and linear probing cell-type annotation. Methods are implemented in biomed-multi-omic ( https://github.com/BiomedSciAI/biomed-multi-omic ), an open-source framework for transcriptomic foundation model development.

转录组基础模型零样本细胞注释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。