arXiv:2601.01089cs.LGq-bio.GN2026-01被引 1

用生物信息流向构建细胞机制模型,实现跨分子层次的统一理解。

Central Dogma Transformer: Towards Mechanism-Oriented AI for Cellular Understanding

  • 按中心法则方向设计注意力机制,整合DNA、RNA、蛋白信息
  • 在K562细胞数据上达到0.503的皮尔逊相关系数,达理论上限的63%
  • 可解释性强,能定位关键调控位点,适合生物机制研究者

解析细胞机制需整合DNA、RNA与蛋白质三类分子信息——这正是分子生物学中心法则所描述的流程。尽管各模态的专用基础模型已取得成功,但彼此孤立,难以建模整体细胞过程。本文提出中心法则变换器(CDT),通过遵循中心法则的方向性逻辑,整合预训练的DNA、RNA和蛋白语言模型。CDT采用定向交叉注意力机制:DNA→RNA注意力模拟转录调控,RNA→蛋白注意力模拟翻译关系,生成融合三者的统一虚拟细胞嵌入。我们以固定(非细胞特异性)的RNA和蛋白嵌入验证CDT v1,在K562细胞CRISPRi增强子扰动数据上取得0.503的皮尔逊相关系数,达到由跨实验变异设定的理论上限0.797的63%。注意力与梯度分析提供互补的可解释视角:案例研究显示二者识别的基因组区域差异显著,梯度分析识别出一个被Hi-C证实同时接触增强子与靶基因的CTCF结合位点。结果表明,符合生物信息流向的AI架构可兼具预测精度与机制可解释性。

原文摘要 · Abstract (English)

Understanding cellular mechanisms requires integrating information across DNA, RNA, and protein - the three molecular systems linked by the Central Dogma of molecular biology. While domain-specific foundation models have achieved success for each modality individually, they remain isolated, limiting our ability to model integrated cellular processes. Here we present the Central Dogma Transformer (CDT), an architecture that integrates pre-trained language models for DNA, RNA, and protein following the directional logic of the Central Dogma. CDT employs directional cross-attention mechanisms - DNA-to-RNA attention models transcriptional regulation, while RNA-to-Protein attention models translational relationships - producing a unified Virtual Cell Embedding that integrates all three modalities. We validate CDT v1 - a proof-of-concept implementation using fixed (non-cell-specific) RNA and protein embeddings - on CRISPRi enhancer perturbation data from K562 cells, achieving a Pearson correlation of 0.503, representing 63% of the theoretical ceiling set by cross-experiment variability (r = 0.797). Attention and gradient analyses provide complementary interpretive windows: in detailed case studies, these approaches highlight largely distinct genomic regions, with gradient analysis identifying a CTCF binding site that Hi-C data showed as physically contacting both enhancer and target gene. These results suggest that AI architectures aligned with biological information flow can achieve both predictive accuracy and mechanistic interpretability.

机制理解多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。