arXiv:2507.05502q-bio.BMcs.LG2025-07ICML被引 14

用折叠能量差预测蛋白结合突变影响,速度快1000倍

Predicting mutational effects on protein binding from folding energy

  • 用预训练折叠模型估算复合物与单体的折叠能量差
  • 在多种数据集上达到与FoldX相当的精度
  • 适合药物设计和结构生物学中的快速突变评估

准确估计突变对蛋白-蛋白结合能的影响是结构生物学和治疗设计中的开放问题。尽管已有多个深度学习方法提出,但因结合数据稀缺,其性能仍不及基于经验力场的计算方法。为此,我们提出一种迁移学习方法,利用蛋白质序列建模和折叠稳定性预测的进展。核心思路是将结合能参数化为复合物折叠能与各组分折叠能之和的差值。我们证明,使用预训练的逆折叠模型作为折叠能代理,可实现强零样本性能,并可通过大量折叠能数据和有限结合能数据进行微调。由此得到的StaB-ddG是首个达到状态最先进力场方法FoldX精度的深度学习预测器,同时速度提升超过1000倍。

原文摘要 · Abstract (English)

Accurate estimation of mutational effects on protein-protein binding energies is an open problem with applications in structural biology and therapeutic design. Several deep learning predictors for this task have been proposed, but, presumably due to the scarcity of binding data, these methods underperform computationally expensive estimates based on empirical force fields. In response, we propose a transfer-learning approach that leverages advances in protein sequence modeling and folding stability prediction for this task. The key idea is to parameterize the binding energy as the difference between the folding energy of the protein complex and the sum of the folding energies of its binding partners. We show that using a pre-trained inverse-folding model as a proxy for folding energy provides strong zero-shot performance, and can be fine-tuned with (1) copious folding energy measurements and (2) more limited binding energy measurements. The resulting predictor, StaB-ddG, is the first deep learning predictor to match the accuracy of the state-of-the-art empirical force-field method FoldX, while offering an over 1,000x speed-up.

蛋白质设计深度学习折叠预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。