arXiv:2412.10411q-bio.QMcs.AI2024-12被引 3

用预训练蛋白语言模型优化密码子,提升疫苗蛋白表达效率

Pre-trained protein language model for codon optimization

  • 用蛋白语言模型提取氨基酸丰富表征,简化密码子优化流程
  • 生成的开放阅读框在稳定性和表达量上优于天然序列和基准序列
  • 特别适合新冠和水痘-带状疱疹病毒疫苗的抗原设计

密码子优化对开放阅读框(ORF)序列的mRNA稳定性与表达效率至关重要,尤其在mRNA疫苗应用中,密码子选择直接影响蛋白产量,进而决定免疫强度。本文探索使用预训练蛋白语言模型(PPLM)获取氨基酸的丰富表征,用于指导密码子优化。该方法将复杂优化问题转化为对PPLM的轻量微调任务。实验表明,所生成的ORF在计算指标上的稳定性和表达性能优于对应天然序列;在针对SARS-CoV-2病毒刺突蛋白和水痘-带状疱疹病毒(VZV)的基准序列对比中也表现更优。结果证明,将PPLM适配于特定抗原的ORF设计具有巨大潜力。

原文摘要 · Abstract (English)

Motivation: Codon optimization of Open Reading Frame (ORF) sequences is essential for enhancing mRNA stability and expression in applications like mRNA vaccines, where codon choice can significantly impact protein yield which directly impacts immune strength. In this work, we investigate the use of a pre-trained protein language model (PPLM) for getting a rich representation of amino acids which could be utilized for codon optimization. This leaves us with a simpler fine-tuning task over PPLM in optimizing ORF sequences. Results: The ORFs generated by our proposed models outperformed their natural counterparts encoding the same proteins on computational metrics for stability and expression. They also demonstrated enhanced performance against the benchmark ORFs used in mRNA vaccines for the SARS-CoV-2 viral spike protein and the varicella-zoster virus (VZV). These results highlight the potential of adapting PPLM for designing ORFs tailored to encode target antigens in mRNA vaccines.

蛋白语言模型密码子优化mRNA疫苗生成设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。