通过掩码基因表达数据,构建高效生物表示学习模型
Effective Biological Representation Learning by Masking Gene Expression

- 采用针对RNA-seq计数数据的掩码自编码方法
- 在140万样本数据集上训练,性能超越百倍更大的数据集模型
- 适合生物医学领域研究者用于基因表达分析
RNA测序产生丰富的基因表达数据,为细胞状态与功能研究提供关键洞察,广泛应用于药物发现。然而,由于技术噪声和实验批次效应,建模此类数据极具挑战性,现有转录组基础模型表现普遍不如线性基线。这引发疑问:深度表示学习是否真能优于直接使用原始计数?本文提出新自监督模型TxFM,聚焦归纳表示学习评估。TxFM采用适配多样RNA-seq计数数据的掩码自编码策略,并通过消融实验确定了实现强迁移性能的关键架构配置。此外,我们构建了公开训练语料库DiverseRNA-1.4M,发现TxFM在此数据上训练所得基因表示,优于在规模大100倍的图谱级语料库上训练的基础模型。结果表明,只要精心结合模型架构与数据整理,归纳式自监督学习是转录组表示学习的可行路径。
原文摘要 · Abstract (English)
RNA sequencing produces rich and diverse datasets of gene expression, offering compelling insights into cellular state and function that have many applications in drug discovery. Modeling such data is challenging due to inherent technical noise and experimental batch effects, as evidenced by many existing transcriptomic foundation models (FMs) underperforming relative to linear baselines. Such results raise the question of whether deep representation learning provides a distinct advantage over the direct use of raw transcript counts. Our work explores this by developing a new self-supervised model, TxFM, with a focus on inductive representation learning evaluations. TxFM employs a masked autoencoding approach tailored to diverse RNA-seq count data, and our ablation study empirically identifies crucial architecture configurations required for strong transfer performance. Additionally, we curate a public training corpus, DiverseRNA-1.4M, and find that TxFM trained on this curated dataset yields high-fidelity gene representations that outperform FMs trained on atlas-scale corpora over 100x larger. Overall, our results indicate that inductive self-supervised learning is a viable modeling approach for transcriptomics representation, provided a careful synthesis of model architecture and training data curation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。