用物理先验提升蛋白互作突变效应预测,效果超越现有方法。
Boltzmann-Aligned Inverse Folding Model as a Predictor of Mutational Effects on Protein-Protein Interactions
- 基于玻尔兹曼分布与逆折叠模型的对齐机制,融合能量与构象信息。
- 无监督与有监督下斯皮尔曼相关系数分别达0.3201和0.5134,优于此前最优结果。
- 适合药物设计中突变影响评估、抗体优化等需精准结合能预测的任务。
预测结合自由能变化(ΔΔG)对于理解与调控蛋白-蛋白相互作用至关重要,这在药物设计中尤为关键。由于实验性ΔΔG数据稀缺,现有方法多聚焦于预训练,忽视了对齐的重要性。本文提出玻尔兹曼对齐技术,将预训练逆折叠模型的知识迁移至ΔΔG预测。通过分析ΔΔG的热力学定义,引入玻尔兹曼分布连接能量与蛋白质构象分布。由于构象分布不可直接计算,我们利用贝叶斯定理,以逆折叠模型提供的对数似然作为ΔΔG估计依据。相比以往基于逆折叠的方法,本方法显式考虑了复合物解离态,在ΔΔG热力学循环中引入物理归纳偏置,实现监督与非监督双重状态最优表现。在SKEMPI v2数据集上的实验表明,无监督与有监督下的斯皮尔曼相关系数分别为0.3201和0.5134,显著优于此前最先进水平(0.2632和0.4324)。此外,我们还展示了该方法在结合能预测、蛋白-蛋白对接及抗体优化任务中的能力。
原文摘要 · Abstract (English)
Predicting the change in binding free energy ($ΔΔG$) is crucial for understanding and modulating protein-protein interactions, which are critical in drug design. Due to the scarcity of experimental $ΔΔG$ data, existing methods focus on pre-training, while neglecting the importance of alignment. In this work, we propose the Boltzmann Alignment technique to transfer knowledge from pre-trained inverse folding models to $ΔΔG$ prediction. We begin by analyzing the thermodynamic definition of $ΔΔG$ and introducing the Boltzmann distribution to connect energy with protein conformational distribution. However, the protein conformational distribution is intractable; therefore, we employ Bayes' theorem to circumvent direct estimation and instead utilize the log-likelihood provided by protein inverse folding models for $ΔΔG$ estimation. Compared to previous inverse folding-based methods, our method explicitly accounts for the unbound state of protein complex in the $ΔΔG$ thermodynamic cycle, introducing a physical inductive bias and achieving both supervised and unsupervised state-of-the-art (SoTA) performance. Experimental results on SKEMPI v2 indicate that our method achieves Spearman coefficients of 0.3201 (unsupervised) and 0.5134 (supervised), significantly surpassing the previously reported SoTA values of 0.2632 and 0.4324, respectively. Futhermore, we demonstrate the capability of our method on binding energy prediction, protein-protein docking and antibody optimization tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。