用微调语言模型+树搜索加速蛋白质计算机演化
Boosting In-Silicon Directed Evolution with Fine-Tuned Protein Language Model and Tree Search
- 微调蛋白语言模型激活进化可塑性
- 测试时用蒙特卡洛树搜索实现高效演化
- 少样本下超越现有方法,适合生物设计
氨基酸突变驱动的蛋白质演化是生命科学的核心。近期蛋白语言模型揭示了丰富的演化模式,为计算机内定向演化带来新可能。但现有方法多依赖启发式策略,尚未有效融合语言模型与强化学习等先进优化技术以自适应学习优演化策略。为此,我们提出AlphaDE框架:首先,利用同源蛋白序列在掩码语言建模下微调预训练蛋白语言模型,激活目标蛋白家族的演化合理性;其次,引入基于蒙特卡洛树搜索的测试时推理,实现由微调模型引导的蛋白质演化。大量基准实验表明,即使采用少样本微调,AlphaDE也显著优于此前最先进方法。案例研究进一步显示,AlphaDE可支持通过计算演化压缩avGFP的蛋白序列空间。
原文摘要 · Abstract (English)
Protein evolution through amino acid mutations is a cornerstone of life sciences. Recent advances in protein language models have shown rich evolutionary patterns, offering unprecedented potential for in-silicon directed evolution. However, existing directed evolution methods largely rely on heuristic evolution strategies and have yet to efficiently integrate the transformative protein language models with advanced optimization techniques, such as reinforcement learning, to adaptively learn superior evolution policies. To bridge this gap, we propose AlphaDE, a novel framework that evolves protein sequences by harnessing the innovative paradigms of large language models, such as fine-tuning and test-time inference. First, AlphaDE fine-tunes pretrained protein language models using masked language modeling on homologous protein sequences to activate the evolutionary plausibility of the interested protein family. Second, AlphaDE introduces test-time inference based on Monte Carlo tree search, which effectively evolves proteins with evolutionary guidance from the fine-tuned protein language model. Extensive benchmark experiments show that AlphaDE remarkably outperforms previous state-of-the-art methods even with few-shot fine-tuning. A case study further demonstrates that AlphaDE supports condensing the protein sequence space of avGFP through computational evolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。