用知识图谱优化蛋白生成,让设计更安全可控
Enhancing Safe and Controllable Protein Generation via Knowledge Preference Optimization
- 引入蛋白安全知识图谱,指导生成过程
- 实验显示有害序列生成率显著降低,功能保持高效
- 适合生物制药、基因工程等高风险场景使用
蛋白质语言模型在序列生成方面展现出强大能力,可有效优化功能并实现从头设计。然而,这些模型也存在生成有害蛋白序列的风险,例如增强病毒传播力或逃避免疫反应,带来重大生物安全与伦理挑战。为此,我们提出一种基于知识引导的偏好优化(KPO)框架,通过蛋白质安全知识图谱整合先验知识。该框架采用高效的图剪枝策略识别优选序列,并利用强化学习最小化生成有害蛋白的风险。实验结果表明,KPO能有效降低有害序列的生成概率,同时保持高水平的功能性,为生物技术领域应用生成模型提供了可靠的安全部署框架。
原文摘要 · Abstract (English)
Protein language models have emerged as powerful tools for sequence generation, offering substantial advantages in functional optimization and denovo design. However, these models also present significant risks of generating harmful protein sequences, such as those that enhance viral transmissibility or evade immune responses. These concerns underscore critical biosafety and ethical challenges. To address these issues, we propose a Knowledge-guided Preference Optimization (KPO) framework that integrates prior knowledge via a Protein Safety Knowledge Graph. This framework utilizes an efficient graph pruning strategy to identify preferred sequences and employs reinforcement learning to minimize the risk of generating harmful proteins. Experimental results demonstrate that KPO effectively reduces the likelihood of producing hazardous sequences while maintaining high functionality, offering a robust safety assurance framework for applying generative models in biotechnology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。