用智能算法加速设计新蛋白,提升药物研发效率。
Deep learning-guided evolutionary optimization for protein design
- 融合遗传算法与贝叶斯优化,高效搜索蛋白序列空间。
- 在肺炎链球菌毒素靶点上快速找到高置信度结合肽。
- 适合蛋白质工程、药物设计人员使用,开源可用。
由于序列空间庞大且序列-功能关系复杂,设计具有特定功能的新蛋白仍具挑战。高效探索该空间以识别满足特定设计目标的序列,对推动治疗和生物技术发展至关重要。本文提出BoGA(贝叶斯优化遗传算法)框架,将进化搜索与贝叶斯优化结合,在代理模型循环中利用遗传算法作为随机候选生成器,基于历史评估结果和代理模型预测优先筛选候选序列,实现数据高效的优化。我们在序列与结构设计任务上进行了基准测试,并应用于设计针对肺炎链球菌关键毒力因子肺炎溶血素的肽类结合剂。BoGA显著加速了高置信度结合剂的发现,展示了在多种设计目标下高效蛋白设计的潜力。该算法集成于BoPep工具套件,代码开源,可于GitHub获取。
原文摘要 · Abstract (English)
Designing novel proteins with desired characteristics remains a significant challenge due to the large sequence space and the complexity of sequence-function relationships. Efficient exploration of this space to identify sequences that meet specific design criteria is crucial for advancing therapeutics and biotechnology. Here, we present BoGA (Bayesian Optimization Genetic Algorithm), a framework that combines evolutionary search with Bayesian optimization to efficiently navigate the sequence space. By integrating a genetic algorithm as a stochastic proposal generator within a surrogate modeling loop, BoGA prioritizes candidates based on prior evaluations and surrogate model predictions, enabling data-efficient optimization. We demonstrate the utility of BoGA through benchmarking on sequence and structure design tasks, followed by its application in designing peptide binders against pneumolysin, a key virulence factor of \textit{Streptococcus pneumoniae}. BoGA accelerates the discovery of high-confidence binders, demonstrating the potential for efficient protein design across diverse objectives. The algorithm is implemented within the BoPep suite and is available under an MIT license at \href{https://github.com/ErikHartman/bopep}{GitHub}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。