arXiv:2409.12995cs.LGcs.AI2024-09

提出新数据划分方法,提升3D结合力模型在小样本下的泛化能力。

Improving generalisability of 3D binding affinity models in low data regimes

  • 构建低相似性数据划分,避免训练测试集泄露,公平对比模型性能。
  • 3D全局模型在小样本下优于蛋白特异性局部模型,且图神经网络性能显著提升。
  • 通过量子数据预训练、分子扩散无监督学习和显式建模氢原子实现突破。

预测蛋白质-配体结合亲和力是计算机辅助药物设计的关键环节。然而,在数据稀缺的情况下,具有泛化能力且表现优异的全局结合亲和力模型仍难以实现。尽管模型架构持续演进,现有基准测试并不适合评估3D结合亲和力模型的泛化性能。此外,如图神经网络(GNN)等3D全局架构的表现未达预期。为解决这些问题,本文提出一种新的PDBBind数据集划分方式,最大程度减少训练集与测试集间的结构相似性泄露,实现不同模型架构之间的公平直接比较。在该低相似性划分下,我们发现:在小样本场景中,3D全局模型总体上优于蛋白特异性局部模型。同时,GNN性能通过三项新贡献得到显著提升:利用量子化学数据进行有监督预训练、通过小分子扩散进行无监督预训练,以及在输入图中显式建模氢原子。我们认为本工作为释放GNN在结合亲和力建模中的潜力提供了新路径。

原文摘要 · Abstract (English)

Predicting protein-ligand binding affinity is an essential part of computer-aided drug design. However, generalisable and performant global binding affinity models remain elusive, particularly in low data regimes. Despite the evolution of model architectures, current benchmarks are not well-suited to probe the generalisability of 3D binding affinity models. Furthermore, 3D global architectures such as GNNs have not lived up to performance expectations. To investigate these issues, we introduce a novel split of the PDBBind dataset, minimizing similarity leakage between train and test sets and allowing for a fair and direct comparison between various model architectures. On this low similarity split, we demonstrate that, in general, 3D global models are superior to protein-specific local models in low data regimes. We also demonstrate that the performance of GNNs benefits from three novel contributions: supervised pre-training via quantum mechanical data, unsupervised pre-training via small molecule diffusion, and explicitly modeling hydrogen atoms in the input graph. We believe that this work introduces promising new approaches to unlock the potential of GNN architectures for binding affinity modelling.

结合亲和力图神经网络小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。