arXiv:2409.17265cs.LGq-bio.QM2024-09被引 3

根据蛋白结构和宿主生物,生成高表达的密码子序列。

CodonMPNN for Organism Specific and Codon Optimal Inverse Folding

  • 基于蛋白骨架和宿主标签,用MPNN生成优化密码子序列。
  • 生成的密码子更接近天然序列,且高适配性序列概率更高。
  • 适合需要高效表达蛋白的工程化设计者使用。

根据蛋白骨架结构生成蛋白质序列是蛋白质工程中的重要技术。在合成工程化蛋白时,通常需将其翻译为DNA并在酵母等宿主中表达。然而,若密码子选择不优,会导致表达效率低下。本文提出CodonMPNN,可根据蛋白骨架结构与宿主生物标签生成优化的密码子序列。当天然DNA序列接近密码子最优时,CodonMPNN能生成比传统启发式方法更高表达量的密码子序列。实验表明,CodonMPNN在保持先前逆折叠方法性能的同时,比基线模型更频繁恢复野生型密码子,并对同一蛋白序列,生成高适配性密码子的概率显著高于低适配性序列。代码已开源于https://github.com/HannesStark/CodonMPNN。

原文摘要 · Abstract (English)

Generating protein sequences conditioned on protein structures is an impactful technique for protein engineering. When synthesizing engineered proteins, they are commonly translated into DNA and expressed in an organism such as yeast. One difficulty in this process is that the expression rates can be low due to suboptimal codon sequences for expressing a protein in a host organism. We propose CodonMPNN, which generates a codon sequence conditioned on a protein backbone structure and an organism label. If naturally occurring DNA sequences are close to codon optimality, CodonMPNN could learn to generate codon sequences with higher expression yields than heuristic codon choices for generated amino acid sequences. Experiments show that CodonMPNN retains the performance of previous inverse folding approaches and recovers wild-type codons more frequently than baselines. Furthermore, CodonMPNN has a higher likelihood of generating high-fitness codon sequences than low-fitness codon sequences for the same protein sequence. Code is available at https://github.com/HannesStark/CodonMPNN.

蛋白质设计密码子优化生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。