arXiv:2411.17795q-bio.QMcs.AI2024-11被引 1

用预训练模型提升低资源酶设计的适应性与泛化能力

Pan-protein Design Learning Enables Task-adaptive Generalization for Low-resource Enzyme Design

  • 通过跨模态对齐序列与结构,迁移预训练知识解决数据少问题
  • 在酶和泛蛋白数据集上表现更优,尤其在域外酶上稳定性强
  • 适合需要快速适配新功能的生物工程场景

计算蛋白质设计(CPD)在生物工程中具有变革潜力,但当前深度学习模型多聚焦通用结构域,难以实现功能特异性设计。本文提出面向功能设计任务的新范式,尤其针对酶这一关键蛋白类别——其常因缺乏特定应用效率而受限。为应对结构数据稀缺问题,提出CrossDesign框架,利用预训练蛋白质语言模型(PPLMs)进行领域自适应。该框架通过将蛋白质结构与序列对齐,将预训练知识迁移至结构模型,克服了结构数据有限的瓶颈。其编码器-解码器架构融合自回归(AR)与非自回归(NAR)状态,在酶数据集及泛蛋白数据集上验证有效。实验表明,CrossDesign在跨域酶设计中性能更优且鲁棒性强;在大规模突变数据上的适应度预测也表现稳定,展现出良好泛化能力。

原文摘要 · Abstract (English)

Computational protein design (CPD) offers transformative potential for bioengineering, but current deep CPD models, focused on universal domains, struggle with function-specific designs. This work introduces a novel CPD paradigm tailored for functional design tasks, particularly for enzymes-a key protein class often lacking specific application efficiency. To address structural data scarcity, we present CrossDesign, a domain-adaptive framework that leverages pretrained protein language models (PPLMs). By aligning protein structures with sequences, CrossDesign transfers pretrained knowledge to structure models, overcoming the limitations of limited structural data. The framework combines autoregressive (AR) and non-autoregressive (NAR) states in its encoder-decoder architecture, applying it to enzyme datasets and pan-proteins. Experimental results highlight CrossDesign's superior performance and robustness, especially with out-of-domain enzymes. Additionally, the model excels in fitness prediction when tested on large-scale mutation data, showcasing its stability.

蛋白质设计酶工程预训练模型低资源学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。