构建蛋白-配体结合亲和力预测的自监督数据集DecoyDB
DecoyDB: A Dataset for Graph Contrastive Learning in Protein-Ligand Binding Affinity Prediction
- 基于高分辨率真实复合物与计算生成的假配体,构建正负样本对
- 预训练模型在PDBbind上表现更优,标签效率提升显著
- 适合药物发现中需高效利用少量标注数据的研究者
预测蛋白-配体复合物的结合亲和力在药物发现中至关重要。然而,高质量标签数据匮乏限制了进展,现有PDBbind数据集仅含不到2万条标注复合物。自监督学习,尤其是图对比学习(GCL),可通过大规模无标签复合物预训练图神经网络模型,再在少量标注数据上微调,突破瓶颈。但该领域面临两大挑战:缺乏包含明确正负样本对的完整无标签数据集,以及需设计适配此类数据特性的GCL算法。为此,我们提出DecoyDB,一个大规模、结构感知的专用数据集,专为蛋白-配体复合物的自监督图对比学习而设计。DecoyDB包含高分辨率真实复合物(<2.5埃)及多种计算生成的假配体结构,其结合姿态从合理到不佳不等(负样本),每个假配体均标注与天然构象的根均方偏差(RMSD)。我们进一步设计定制化GCL框架,基于DecoyDB预训练图神经网络,并在PDBbind标签上微调。大量实验表明,经DecoyDB预训练的模型在准确率、标签效率和泛化能力上均表现优异。
原文摘要 · Abstract (English)
Predicting the binding affinity of protein-ligand complexes plays a vital role in drug discovery. Unfortunately, progress has been hindered by the lack of large-scale and high-quality binding affinity labels. The widely used PDBbind dataset has fewer than 20K labeled complexes. Self-supervised learning, especially graph contrastive learning (GCL), provides a unique opportunity to break the barrier by pre-training graph neural network models based on vast unlabeled complexes and fine-tuning the models on much fewer labeled complexes. However, the problem faces unique challenges, including a lack of a comprehensive unlabeled dataset with well-defined positive/negative complex pairs and the need to design GCL algorithms that incorporate the unique characteristics of such data. To fill the gap, we propose DecoyDB, a large-scale, structure-aware dataset specifically designed for self-supervised GCL on protein-ligand complexes. DecoyDB consists of high-resolution ground truth complexes (less than 2.5 Angstrom) and diverse decoy structures with computationally generated binding poses that range from realistic to suboptimal (negative pairs). Each decoy is annotated with a Root Mean Squared Deviation (RMSD) from the native pose. We further design a customized GCL framework to pre-train graph neural networks based on DecoyDB and fine-tune the models with labels from PDBbind. Extensive experiments confirm that models pre-trained with DecoyDB achieve superior accuracy, label efficiency, and generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。