大模型暗中具备优化蛋白序列的能力,可高效设计高适应性蛋白。
Large Language Model is Secretly a Protein Sequence Optimizer
- 用类定向进化方法让大模型在约束条件下优化蛋白序列
- 在合成与真实数据集上均实现高适应性序列生成
- 适合生物设计、药物开发等需要快速迭代的场景
我们研究蛋白序列工程问题,目标是从野生型序列出发,寻找高适应性蛋白序列。定向进化是该领域的主流范式,通过迭代生成变异体并依赖实验反馈进行筛选。本文发现,尽管大语言模型(LLMs)仅在海量文本上训练,但其本质上可作为蛋白序列优化器。通过结合定向进化方法,LLM 能在帕累托最优和实验预算双重约束下完成优化,在合成与真实适应性景观上均取得成功。
原文摘要 · Abstract (English)
We consider the protein sequence engineering problem, which aims to find protein sequences with high fitness levels, starting from a given wild-type sequence. Directed evolution has been a dominating paradigm in this field which has an iterative process to generate variants and select via experimental feedback. We demonstrate large language models (LLMs), despite being trained on massive texts, are secretly protein sequence optimizers. With a directed evolutionary method, LLM can perform protein engineering through Pareto and experiment-budget constrained optimization, demonstrating success on both synthetic and experimental fitness landscapes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。