用任意文本描述生成合理蛋白序列,性能领先
ProtDAT: A Unified Framework for Protein Sequence Design from Any Protein Text Description
- 将蛋白序列与文本通过跨模态注意力融合设计
- 在瑞士蛋白数据库上提升pLDDT6%、TM-score0.26
- 适合需要从文字描述生成蛋白的研究者
蛋白质设计在药物开发和酶工程等领域具有重要应用前景。然而,仅依赖大规模语言模型预训练与微调的方法难以捕捉多模态蛋白质数据间的关系。为此,我们提出ProtDAT,一种从任意蛋白文本描述生成新蛋白序列的端到端细粒度框架。该框架基于蛋白质数据特性,将序列与文本视为统一整体而非独立实体,采用创新的多模态交叉注意力机制,在基础层面实现无缝融合。实验表明,ProtDAT在蛋白序列生成任务中达到当前最优表现,显著提升合理性、功能性和结构相似性。在包含20,000个文本-序列对的Swiss-Prot数据集上,其pLDDT提升6%,TM-score提高0.26,均方根偏差(RMSD)降低1.2 Å,展现出强大的蛋白设计潜力。
原文摘要 · Abstract (English)
Protein design has become a critical method in advancing significant potential for various applications such as drug development and enzyme engineering. However, protein design methods utilizing large language models with solely pretraining and fine-tuning struggle to capture relationships in multi-modal protein data. To address this, we propose ProtDAT, a de novo fine-grained framework capable of designing proteins from any descriptive protein text input. ProtDAT builds upon the inherent characteristics of protein data to unify sequences and text as a cohesive whole rather than separate entities. It leverages an innovative multi-modal cross-attention, integrating protein sequences and textual information for a foundational level and seamless integration. Experimental results demonstrate that ProtDAT achieves the state-of-the-art performance in protein sequence generation, excelling in rationality, functionality, structural similarity, and validity. On 20,000 text-sequence pairs from Swiss-Prot, it improves pLDDT by 6%, TM-score by 0.26, and reduces RMSD by 1.2 Å, highlighting its potential to advance protein design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。