用大模型优化实现精准可控的蛋白质序列设计
Controllable Protein Sequence Generation with LLM Preference Optimization
- 通过多列表偏好优化微调蛋白大模型,提升生成可控性
- 在单属性与多属性生成任务中均达当前最优性能
- 适合需要精准设计功能蛋白的研究者使用
设计具有特定属性的蛋白质为解决生物医学挑战提供了重要途径。预训练的蛋白质大语言模型(LLMs)在蛋白质序列生成方面已展现出良好效果。然而,现有方法在控制序列生成以满足特定属性时,仍存在功能性和结构稳定性不足的问题。本文提出一种名为CtrlProt的新颖可控制蛋白质设计方法。通过引入一种新的多列表偏好优化策略对蛋白大模型进行微调,显著提升了生成质量,并支持多属性可控生成。实验表明,CtrlProt能有效满足功能性和结构稳定性要求,在单属性和多属性蛋白质序列生成任务中均达到当前最优水平。
原文摘要 · Abstract (English)
Designing proteins with specific attributes offers an important solution to address biomedical challenges. Pre-trained protein large language models (LLMs) have shown promising results on protein sequence generation. However, to control sequence generation for specific attributes, existing work still exhibits poor functionality and structural stability. In this paper, we propose a novel controllable protein design method called CtrlProt. We finetune a protein LLM with a new multi-listwise preference optimization strategy to improve generation quality and support multi-attribute controllable generation. Experiments demonstrate that CtrlProt can meet functionality and structural stability requirements effectively, achieving state-of-the-art performance in both single-attribute and multi-attribute protein sequence generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。