用Transformer学习基因表达特征,提升癌症分类与预后预测能力
Transformer-Based Representation Learning for Robust Gene Expression Modeling and Cancer Prognosis
- 基于掩码重建训练基因表达数据,捕捉上千基因共表达关系
- 在少量基因下实现最优分类,显著改善生存预测与缺失值填补
- 注意力机制揭示跨癌种的生物意义基因模式,适合临床数据不全场景
Transformer模型在自然语言和视觉任务中表现卓越,但其在基因表达分析中的应用受限于数据稀疏、高维及缺失值问题。本文提出GexBERT,一种基于Transformer的自编码框架,通过大规模转录组数据预训练,采用掩码与恢复目标学习上下文感知的基因嵌入,捕捉数千基因间的共表达关系。我们在三项关键癌症研究任务中评估GexBERT:泛癌分类、癌症特异性生存预测和缺失值插补。结果表明,GexBERT在有限基因子集下达到最先进的分类准确率,通过恢复预后锚基因表达提升了生存预测性能,并在高缺失率下优于传统插补方法。此外,其基于注意力的可解释性揭示了跨癌种的生物学意义基因模式。这些发现证明GexBERT是一种可扩展且高效的基因表达建模工具,在基因覆盖不全或不完整场景中具有转化潜力。
原文摘要 · Abstract (English)
Transformer-based models have achieved remarkable success in natural language and vision tasks, but their application to gene expression analysis remains limited due to data sparsity, high dimensionality, and missing values. We present GexBERT, a transformer-based autoencoder framework for robust representation learning of gene expression data. GexBERT learns context-aware gene embeddings by pretraining on large-scale transcriptomic profiles with a masking and restoration objective that captures co-expression relationships among thousands of genes. We evaluate GexBERT across three critical tasks in cancer research: pan-cancer classification, cancer-specific survival prediction, and missing value imputation. GexBERT achieves state-of-the-art classification accuracy from limited gene subsets, improves survival prediction by restoring expression of prognostic anchor genes, and outperforms conventional imputation methods under high missingness. Furthermore, its attention-based interpretability reveals biologically meaningful gene patterns across cancer types. These findings demonstrate the utility of GexBERT as a scalable and effective tool for gene expression modeling, with translational potential in settings where gene coverage is limited or incomplete.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。