arXiv:2505.04823cs.LGq-bio.BM2025-05被引 6

无需重新训练,即可实时引导蛋白质生成模型设计具有特定性质的序列。

ProteinGuide: On-the-fly property guidance for protein sequence generative models

  • 提出统一框架,实现对多种生成模型的即时属性引导。
  • 单轮实验即提升腺嘌呤碱基编辑器活性,超越七轮定向进化效果。
  • 适用于稳定性与活性等冲突属性的协同优化,适合实验设计者使用。

序列生成模型正在重塑蛋白质工程。然而,目前尚无一种可泛化的方法,在不重新训练生成模型的前提下,对辅助信息(如实验数据)进行条件建模。本文提出ProteinGuide,一种“即时”条件化方法,适用于包括掩码语言模型(如ESM3)、任意顺序自回归模型(如ProteinMPNN)以及扩散和流匹配模型(如MultiFlow)在内的广泛模型类别。该方法基于我们对这些模型类别的统一统计视角。作为验证,我们进行了多项体外实验:首先引导预训练模型生成具有指定性质(如更高稳定性和活性)的蛋白质;其次设计同时优化两个相互冲突性质的蛋白质;最后在湿实验中应用该方法,仅用一个包含2000个变体的池化文库数据,成功提升了体内腺嘌呤碱基编辑器的编辑活性。结果显示,单轮ProteinGuide的效果优于以往七轮定向进化所达到的水平。

原文摘要 · Abstract (English)

Sequence generative models are transforming protein engineering. However, no principled framework exists for conditioning these models on auxiliary information, such as experimental data, without additional training of a generative model. Herein, we present ProteinGuide, a method for such "on-the-fly" conditioning, amenable to a broad class of protein generative models including Masked Language Models (e.g. ESM3), any-order auto-regressive models (e.g. ProteinMPNN) as well as diffusion and flow matching models (e.g. MultiFlow). ProteinGuide stems from our unifying view of these model classes under a single statistical framework. As proof of principle, we perform several in silico experiments. We first guide pre-trained generative models to design proteins with user-specified properties, such as higher stability or activity. Next, we design for optimizing two desired properties that are in tension with each other. Finally, we apply our method in the wet lab, using ProteinGuide to increase the editing activity of an adenine base editor in vivo with data from only a single pooled library of 2,000 variants. We find that a single round of ProteinGuide achieves a higher editing efficiency than was previously achieved using seven rounds of directed evolution.

蛋白质生成属性引导定向进化生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。