arXiv:2502.07299cs.LGcs.AI2025-02被引 4

用中心法则统一多组学数据,实现基因到蛋白的全流程建模

Life-Code: Central Dogma Modeling with Multi-Omics Sequence Unification

  • 将RNA、蛋白质逆向转为核苷酸序列,构建统一输入格式
  • 通过掩码建模预训练,捕捉编码与非编码区的复杂交互
  • 融合现有蛋白语言模型知识,精准预测蛋白质结构

DNA、RNA与蛋白质之间的相互作用是生物过程的核心,体现了分子生物学的中心法则。尽管现代生物预训练模型在独立分析这些大分子方面取得显著进展,但它们的内在关联仍研究不足。本文基于中心法则重构数据与模型流程,提出一个覆盖多种生物功能的综合性框架Life-Code。在数据层面,提出统一管道,通过逆转录RNA和反向翻译氨基酸为基于核苷酸的序列,实现多组学数据整合。在模型层面,设计密码子分词器和混合长序列架构,利用掩码建模预训练捕捉编码区与非编码区间的交互。为建模翻译与折叠过程,Life-Code通过知识蒸馏从现成蛋白语言模型中学习对应氨基酸的蛋白质结构。该设计使Life-Code能够捕获遗传序列内的复杂交互,更全面地理解多组学数据。大量实验表明,Life-Code在三个组学的多项任务上达到当前最优表现,展现了其推动多组学分析与解读的巨大潜力。

原文摘要 · Abstract (English)

The interactions between DNA, RNA, and proteins are fundamental to biological processes, as illustrated by the central dogma of molecular biology. Although modern biological pre-trained models have achieved great success in analyzing these macromolecules individually, their interconnected nature remains underexplored. This paper follows the guidance of the central dogma to redesign both the data and model pipeline and offers a comprehensive framework, Life-Code, that spans different biological functions. As for data flow, we propose a unified pipeline to integrate multi-omics data by reverse-transcribing RNA and reverse-translating amino acids into nucleotide-based sequences. As for the model, we design a codon tokenizer and a hybrid long-sequence architecture to encode the interactions between coding and non-coding regions through masked modeling pre-training. To model the translation and folding process with coding sequences, Life-Code learns protein structures of the corresponding amino acids by knowledge distillation from off-the-shelf protein language models. Such designs enable Life-Code to capture complex interactions within genetic sequences, providing a more comprehensive understanding of multi-omics with the central dogma. Extensive experiments show that Life-Code achieves state-of-the-art results on various tasks across three omics, highlighting its potential for advancing multi-omics analysis and interpretation.

多组学中心法则序列建模蛋白结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。