arXiv:2507.03318cs.LGcs.AI2025-07被引 1

用图神经网络和正则化方法,提升靶点特异性药物结合力预测的可解释性。

Structure-Aware Compound-Protein Affinity Prediction via Graph Neural Networks with Group Lasso Regularization

  • 基于活性悬崖对的结构与性质信息,构建可解释的图神经网络模型。
  • 在6个酪氨酸激酶靶点上,预测误差降低,相关系数提升。
  • 通过稀疏组套索正则化,精准定位影响结合力的关键分子片段。

可解释人工智能加速药物发现,通过改进分子表征学习、识别关键分子结构并合理化药物性质预测。然而,针对特定靶点的结构-活性关系建模仍面临挑战,因化合物-蛋白相互作用数据有限,且化学取代基或局部结构微小变化可能引起性质显著差异。为此,我们提出一种图神经网络(GNN)框架,利用靶向特定蛋白的活性悬崖分子对的性质与结构信息,预测化合物-蛋白亲和力(以半最大抑制浓度IC50衡量),并解释性质差异。通过引入结构感知损失函数及组套索、稀疏组套索正则化,实现对与活性差异相关的分子子图的剪枝与突出。该框架应用于Src、Abl、Tec家族及ALK激酶的6个靶点,整合共性与非共性节点信息,结合稀疏组套索正则化,显著降低根均方误差,提高皮尔逊相关系数。正则化还提升了特征归因效果,增强图级全局方向得分与原子级着色准确率,支持更可解释的药物发现流程,尤其适用于先导化合物优化中识别关键分子亚结构。

原文摘要 · Abstract (English)

Explainable artificial intelligence approaches accelerate drug discovery by improving molecular representation learning, identifying key molecular structures, and rationalizing drug property prediction. However, developing end-to-end explainable models for target-specific structure-activity relationship modeling remains challenging because compound-protein interaction data are often limited for individual targets, and small changes in chemical substituents or local structural motifs can cause large differences in molecular properties. Therefore, effectively leveraging structural and property information to identify key moieties associated with compound-protein affinity is essential. We propose a graph neural network (GNN) framework that uses property and structural information from activity-cliff molecule pairs targeting specific proteins to predict compound-protein affinity, measured by half-maximal inhibitory concentration (IC50), and explain property differences. To improve explainability, we trained GNNs with structure-aware loss functions using group lasso and sparse group lasso regularization, which prune and highlight molecular subgraphs relevant to activity differences. We applied this framework to activity-cliff data from molecules targeting six tyrosine-protein kinases across the Src, Abl, and Tec families, as well as anaplastic lymphoma kinase. Integrating common- and uncommon-node information with sparse group lasso improved target-specific molecular property prediction, producing lower root mean square errors and higher Pearson correlation coefficients. Regularization also enhanced GNN feature attribution by improving graph-level global direction scores and atom-level coloring accuracy. These results support more interpretable drug discovery pipelines, particularly for identifying critical molecular substructures during lead optimization.

药物发现图神经网络可解释性分子表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。