模型首次打通从DNA到蛋白的全流程可解释预测,揭示了转录与翻译的反向调控规律。
Central Dogma Transformer III: Interpretable AI Across DNA, RNA, and Protein
- 分核质两阶段建模,模拟细胞内转录与翻译的空间过程。
- 在5个基因上蛋白质预测相关系数达0.969,显著提升上游表征质量。
- 可无实验数据筛查全基因组副作用,适合药物靶点评估与机制研究。
生物AI模型日益能预测复杂的细胞响应,但其学习表征仍与真实分子过程脱节。我们提出CDT-III,首次将机制导向的AI扩展至整个中心法则:DNA、RNA和蛋白。其双阶段虚拟细胞嵌入器架构模仿细胞空间分隔:VCE-N建模核内转录,VCE-C建模胞质翻译。在五个保留基因上,CDT-III实现每基因RNA相关系数r=0.843,蛋白r=0.969。加入蛋白预测任务后,RNA性能从r=0.804提升至0.843,表明下游任务可正则化上游表示。蛋白监督使DNA层面可解释性增强,CTCF富集度提升30%。分析实测mRNA与蛋白响应发现,多数有可检测mRNA变化的基因在蛋白水平呈现相反变化(|log2FC|>0.01时占66.7%,|log2FC|>0.02时升至87.5%),暴露了仅依赖RNA扰动模型的根本局限。尽管存在普遍的方向不一致,CDT-III仍能正确预测mRNA与蛋白响应。应用于模拟CD52敲低以类比阿仑单抗,模型准确预测29/29个蛋白变化,并无临床数据复现5/7已知临床副作用。基于梯度的副作用分析仅需未扰动基线数据(r=0.939),实现对全部2,361个基因的无实验筛查。
原文摘要 · Abstract (English)
Biological AI models increasingly predict complex cellular responses, yet their learned representations remain disconnected from the molecular processes they aim to capture. We present CDT-III, which extends mechanism-oriented AI across the full central dogma: DNA, RNA, and protein. Its two-stage Virtual Cell Embedder architecture mirrors the spatial compartmentalization of the cell: VCE-N models transcription in the nucleus and VCE-C models translation in the cytosol. On five held-out genes, CDT-III achieves per-gene RNA r=0.843 and protein r=0.969. Adding protein prediction improves RNA performance (r=0.804 to 0.843), demonstrating that downstream tasks regularize upstream representations. Protein supervision sharpens DNA-level interpretability, increasing CTCF enrichment by 30%. Analysis of experimentally measured mRNA and protein responses reveals that the majority of genes with observable mRNA changes show opposite protein-level changes (66.7% at |log2FC|>0.01, rising to 87.5% at |log2FC|>0.02), exposing a fundamental limitation of RNA-only perturbation models. Despite this pervasive direction discordance, CDT-III correctly predicts both mRNA and protein responses. Applied to in silico CD52 knockdown approximating Alemtuzumab, the model predicts 29/29 protein changes correctly and rediscovers 5 of 7 known clinical side effects without clinical data. Gradient-based side effect profiling requires only unperturbed baseline data (r=0.939), enabling screening of all 2,361 genes without new experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。