arXiv:2505.18966cs.LGcs.AI2025-05NeurIPS被引 12

用自然蛋白片段提升生成蛋白的可折叠性

Protein Design with Dynamic Protein Vocabulary

  • 从自然蛋白中动态检索片段,增强生成序列的结构合理性
  • 仅用0.04%数据达到相近功能匹配,可折叠蛋白比例提升9.6%
  • 适合生物设计、生成模型研究者快速构建高可折叠蛋白

蛋白质设计是生物技术中的核心挑战,旨在从巨大可能空间中设计出具有特定功能的新序列。尽管深度生成模型已能根据文本描述实现功能导向设计,但在结构合理性方面仍存在不足。受经典蛋白质设计方法启发,我们探索将自然蛋白质片段引入生成模型是否能提升可折叠性。实验表明,即使随机引入片段也能改善折叠性能。基于此,我们提出ProDVa:结合文本编码器、蛋白质语言模型和片段编码器,根据功能描述动态检索并整合蛋白片段。实验显示,该方法在功能对齐上与顶尖模型相当,训练数据使用量不足0.04%,同时显著提升可折叠性——pLDDT高于70的蛋白比例提升7.38%,PAE低于10的比例提升9.6%。

原文摘要 · Abstract (English)

Protein design is a fundamental challenge in biotechnology, aiming to design novel sequences with specific functions within the vast space of possible proteins. Recent advances in deep generative models have enabled function-based protein design from textual descriptions, yet struggle with structural plausibility. Inspired by classical protein design methods that leverage natural protein structures, we explore whether incorporating fragments from natural proteins can enhance foldability in generative models. Our empirical results show that even random incorporation of fragments improves foldability. Building on this insight, we introduce ProDVa, a novel protein design approach that integrates a text encoder for functional descriptions, a protein language model for designing proteins, and a fragment encoder to dynamically retrieve protein fragments based on textual functional descriptions. Experimental results demonstrate that our approach effectively designs protein sequences that are both functionally aligned and structurally plausible. Compared to state-of-the-art models, ProDVa achieves comparable function alignment using less than 0.04% of the training data, while designing significantly more well-folded proteins, with the proportion of proteins having pLDDT above 70 increasing by 7.38% and those with PAE below 10 increasing by 9.6%.

蛋白质设计生成模型结构可折叠性片段检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。