用图注意力网络精准预测蛋白结合位点,兼顾效率与可解释性
Edge-aware GAT-based protein binding site prediction
- 构建原子级图结构,融合几何、二级结构等多维特征
- 在基准数据集上达到0.93的ROC-AUC,优于现有方法
- 适合药物设计与蛋白质互作研究者使用
准确识别蛋白结合位点对理解生物分子相互作用机制及合理设计药物靶点至关重要。传统方法在捕捉复杂空间构象时难以兼顾预测精度与计算效率。为此,我们提出一种边缘感知图注意力网络(Edge-aware GAT),用于精细预测包括蛋白质、核酸、离子、配体和脂质在内的多种生物分子的结合位点。该方法构建原子级图,整合几何描述符、DSSP衍生二级结构及相对溶剂暴露度(RSA)等多维结构特征,生成具有空间感知能力的嵌入向量。通过在注意力机制中引入原子间距离与方向向量作为边特征,显著提升模型表征能力。在基准数据集上,该模型在蛋白-蛋白结合位点预测中取得0.93的ROC-AUC,优于多个前沿方法。方向张量传播与残基级注意力池化进一步提升了结合位点定位精度与局部结构细节捕捉能力。PyMOL可视化验证了模型的实用性和可解释性。为促进社区应用,我们已公开部署在线服务器(http://119.45.201.89:5000/)。本方法为蛋白功能位点识别提供了高效、通用且可解释的新方案。
原文摘要 · Abstract (English)
Accurate identification of protein binding sites is crucial for understanding biomolecular interaction mechanisms and for the rational design of drug targets. Traditional predictive methods often struggle to balance prediction accuracy with computational efficiency when capturing complex spatial conformations. To address this challenge, we propose an Edge-aware Graph Attention Network (Edge-aware GAT) model for the fine-grained prediction of binding sites across various biomolecules, including proteins, DNA/RNA, ions, ligands, and lipids. Our method constructs atom-level graphs and integrates multidimensional structural features, including geometric descriptors, DSSP-derived secondary structure, and relative solvent accessibility (RSA), to generate spatially aware embedding vectors. By incorporating interatomic distances and directional vectors as edge features within the attention mechanism, the model significantly enhances its representation capacity. On benchmark datasets, our model achieves an ROC-AUC of 0.93 for protein-protein binding site prediction, outperforming several state-of-the-art methods. The use of directional tensor propagation and residue-level attention pooling further improves both binding site localization and the capture of local structural details. Visualizations using PyMOL confirm the model's practical utility and interpretability. To facilitate community access and application, we have deployed a publicly accessible web server at http://119.45.201.89:5000/. In summary, our approach offers a novel and efficient solution that balances prediction accuracy, generalization, and interpretability for identifying functional sites in proteins.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。