arXiv:2510.24053cs.LGq-bio.QM2025-10被引 1

用智能选突变+语言模型补数据,让少样本蛋白优化更高效

Low-N Protein Activity Optimization with FolDE

  • 用蛋白语言模型生成自然序列,弥补实验数据不足
  • 在20个蛋白上比最好基线多发现23%顶级突变
  • 适合资源有限的实验室做蛋白质功能优化

传统蛋白质优化需大量突变体构建与测试,成本高昂。主动学习辅助定向进化(ALDE)通过预测最优突变并迭代验证来降低成本。但现有方法在每轮选择最高预测值突变,导致训练数据同质化,影响后续预测准确性。本文提出FolDE,一种旨在最大化终局成功率的ALDE方法。在20个蛋白目标的模拟实验中,FolDE发现的前10%优质突变比最佳基线多23%(p=0.005),找到前1%优质突变的概率高55%。其核心在于基于自然性的冷启动策略,利用蛋白语言模型输出补充有限活性测量数据,提升预测精度。还引入恒谎批次选择器以增强批次多样性,虽在基准测试中效果有限,但在多突变场景中意义重要。完整工作流程开源,使高效蛋白优化惠及所有实验室。

原文摘要 · Abstract (English)

Proteins are traditionally optimized through the costly construction and measurement of many mutants. Active Learning-assisted Directed Evolution (ALDE) alleviates that cost by predicting the best improvements and iteratively testing mutants to inform predictions. However, existing ALDE methods face a critical limitation: selecting the highest-predicted mutants in each round yields homogeneous training data insufficient for accurate prediction models in subsequent rounds. Here we present FolDE, an ALDE method designed to maximize end-of-campaign success. In simulations across 20 protein targets, FolDE discovers 23% more top 10% mutants than the best baseline ALDE method (p=0.005) and is 55% more likely to find top 1% mutants. FolDE achieves this primarily through naturalness-based warm-starting, which augments limited activity measurements with protein language model outputs to improve activity prediction. We also introduce a constant-liar batch selector, which improves batch diversity; this is important in multi-mutation campaigns but had limited effect in our benchmarks. The complete workflow is freely available as open-source software, making efficient protein optimization accessible to any laboratory.

蛋白优化主动学习语言模型定向进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。