用扩散机制联合学习结构与序列信息,提升酶稳定性预测精度
DGTN: Graph-Enhanced Transformer with Diffusive Attention Gating Mechanism for Enzyme DDG Prediction
- 通过双向扩散机制融合图网络与注意力机制
- 在ProTherm和SKEMPI上达0.87相关性,误差1.21 kcal/mol
- 理论证明收敛速度为O(1/sqrt(T)),适合蛋白工程研究
预测氨基酸突变对酶热力学稳定性(DDG)的影响是蛋白质工程和药物设计的基础。尽管深度学习方法已展现潜力,但通常独立处理序列与结构信息,难以捕捉局部几何结构与全局序列模式之间的复杂耦合。我们提出DGTN(扩散图-变换器网络),一种新架构,通过扩散机制共同学习图神经网络(GNN)的结构先验权重与变换器注意力。核心创新在于双向扩散过程:(1) 由GNN生成的结构嵌入通过可学习扩散核引导变换器注意力;(2) 变换器表示通过注意力调制的图更新优化GNN消息传递。我们提供严格的数学分析,证明该联合学习方案优于独立处理。在ProTherm和SKEMPI基准上,DGTN达到领先性能(皮尔逊相关系数Rho = 0.87,RMSE = 1.21 kcal/mol),较最优基线提升6.2%。消融实验表明,扩散机制贡献4.8点相关性提升。理论分析证明扩散注意力收敛至最优结构-序列耦合,收敛速率O(1/sqrt(T)),其中T为扩散步数。本工作建立了一个通过可学习扩散整合异构蛋白表示的原理性框架。
原文摘要 · Abstract (English)
Predicting the effect of amino acid mutations on enzyme thermodynamic stability (DDG) is fundamental to protein engineering and drug design. While recent deep learning approaches have shown promise, they often process sequence and structure information independently, failing to capture the intricate coupling between local structural geometry and global sequential patterns. We present DGTN (Diffused Graph-Transformer Network), a novel architecture that co-learns graph neural network (GNN) weights for structural priors and transformer attention through a diffusion mechanism. Our key innovation is a bidirectional diffusion process where: (1) GNN-derived structural embeddings guide transformer attention via learnable diffusion kernels, and (2) transformer representations refine GNN message passing through attention-modulated graph updates. We provide rigorous mathematical analysis showing this co-learning scheme achieves provably better approximation bounds than independent processing. On ProTherm and SKEMPI benchmarks, DGTN achieves state-of-the-art performance (Pearson Rho = 0.87, RMSE = 1.21 kcal/mol), with 6.2% improvement over best baselines. Ablation studies confirm the diffusion mechanism contributes 4.8 points to correlation. Our theoretical analysis proves the diffused attention converges to optimal structure-sequence coupling, with convergence rate O(1/sqrt(T) ) where T is diffusion steps. This work establishes a principled framework for integrating heterogeneous protein representations through learnable diffusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。