arXiv:2505.22494cs.LG2025-05NeurIPS被引 2

用主动学习设计高适配度且新颖的蛋白质序列,突破野生型附近限制。

ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods

  • 用预训练生成模型配合可更新代理模型,结合残基选择与生物约束采样。
  • 在代理模型不准确时仍有效,能稳定生成高适配度且新颖的序列。
  • 适合需要探索新蛋白空间的高效工程任务,如药物设计或酶改造。

在数据效率有限的蛋白质工程中,设计兼具高适配度与新颖性的蛋白序列是一项挑战。超越野生型邻域的探索常导致生物学上不可行的序列,或依赖在新区域失去保真的代理模型。本文提出 ProSpero,一种主动学习框架:冻结的预训练生成模型由从真实反馈(oracle)更新的代理模型引导。通过整合与适配度相关的残基选择和生物约束的顺序蒙特卡洛采样,该方法可在保持生物学合理性的同时,拓展至野生型邻域之外。我们证明,即使代理模型存在偏差,该框架依然有效。ProSpero 在多种蛋白质工程任务中持续优于或匹配现有方法,成功检索到兼具高适配度与新颖性的序列。

原文摘要 · Abstract (English)

Designing protein sequences of both high fitness and novelty is a challenging task in data-efficient protein engineering. Exploration beyond wild-type neighborhoods often leads to biologically implausible sequences or relies on surrogate models that lose fidelity in novel regions. Here, we propose ProSpero, an active learning framework in which a frozen pre-trained generative model is guided by a surrogate updated from oracle feedback. By integrating fitness-relevant residue selection with biologically-constrained Sequential Monte Carlo sampling, our approach enables exploration beyond wild-type neighborhoods while preserving biological plausibility. We show that our framework remains effective even when the surrogate is misspecified. ProSpero consistently outperforms or matches existing methods across diverse protein engineering tasks, retrieving sequences of both high fitness and novelty.

蛋白质设计主动学习生成模型生物约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。