arXiv:2506.13196cs.LG2025-06

融合生化知识的深度学习模型,提升药物靶点结合亲和力预测准确率

KEPLA: A Knowledge-Enhanced Deep Learning Framework for Accurate Protein-Ligand Binding Affinity Prediction

  • 引入基因本体与配体属性知识增强模型表征
  • 在两个数据集上优于现有最佳方法,跨领域表现稳定
  • 可解释性分析揭示关键作用机制,适合药物研发人员

准确预测蛋白质-配体结合亲和力对药物发现至关重要。尽管近期深度学习方法已取得良好效果,但多仅依赖蛋白与配体的结构特征,忽视与其结合亲和力相关的生化知识。为此,我们提出KEPLA——一种新型深度学习框架,显式整合基因本体(Gene Ontology)和配体性质等先验知识以提升预测性能。KEPLA以蛋白序列和配体分子图作为输入,优化两个互补目标:(1) 对齐全局表征与知识图谱关系,捕捉领域特异性生化洞察;(2) 利用局部表征间的交叉注意力构建细粒度联合嵌入用于预测。在两个基准数据集上的实验表明,无论是在域内还是跨域场景下,KEPLA均持续优于当前最优基线。此外,基于知识图谱关系和交叉注意力图的可解释性分析,为模型预测机制提供了重要洞察。

原文摘要 · Abstract (English)

Accurate prediction of protein-ligand binding affinity is critical for drug discovery. While recent deep learning approaches have demonstrated promising results, they often rely solely on structural features of proteins and ligands, overlooking their valuable biochemical knowledge associated with binding affinity. To address this limitation, we propose KEPLA, a novel deep learning framework that explicitly integrates prior knowledge from Gene Ontology and ligand properties to enhance prediction performance. KEPLA takes protein sequences and ligand molecular graphs as input and optimizes two complementary objectives: (1) aligning global representations with knowledge graph relations to capture domain-specific biochemical insights, and (2) leveraging cross attention between local representations to construct fine-grained joint embeddings for prediction. Experiments on two benchmark datasets across both in-domain and cross-domain scenarios demonstrate that KEPLA consistently outperforms state-of-the-art baselines. Furthermore, interpretability analyses based on knowledge graph relations and cross attention maps provide valuable insights into the underlying predictive mechanisms.

蛋白质-配体深度学习药物发现可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。