用激活值调控让蛋白语言模型精准生成目标功能蛋白
Steering Protein Language Models
- 通过编辑神经网络激活值引导蛋白序列生成
- 在溶菌酶类序列生成中实现功能定向优化
- 无需额外训练,适配多种蛋白模型架构
蛋白语言模型(PLMs)基于海量天然蛋白的进化数据预训练,已成为蛋白质设计的重要工具。然而,由于输出控制困难,现有PLMs难以生成具有精确功能或性质的蛋白。本文探索了源自大语言模型文本生成控制的激活值调节技术在PLM中的应用,提出一种简单有效的激活编辑方法,并引入新型编辑位点识别模块以支持蛋白优化。在溶菌酶样序列生成与优化任务上,实验表明该方法可无缝集成到自编码和自回归型PLM中,且无需额外训练。结果验证了利用基础模型实现精准蛋白工程的可行性。
原文摘要 · Abstract (English)
Protein Language Models (PLMs), pre-trained on extensive evolutionary data from natural proteins, have emerged as indispensable tools for protein design. While powerful, PLMs often struggle to produce proteins with precisely specified functionalities or properties due to inherent challenges in controlling their outputs. In this work, we investigate the potential of Activation Steering, a technique originally developed for controlling text generation in Large Language Models (LLMs), to direct PLMs toward generating protein sequences with targeted properties. We propose a simple yet effective method that employs activation editing to steer PLM outputs, and extend this approach to protein optimization through a novel editing site identification module. Through comprehensive experiments on lysozyme-like sequence generation and optimization, we demonstrate that our methods can be seamlessly integrated into both auto-encoding and autoregressive PLMs without requiring additional training. These results highlight a promising direction for precise protein engineering using foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。