构建含蛋白修饰的完整DAVIS数据集,推动精准药物亲和力预测
Towards Precision Protein-Ligand Affinity Prediction Benchmark: A Complete and Modification-Aware DAVIS Dataset
- 构建含4032个激酶-配体对的修饰感知数据集
- 发现基于对接模型在无训练场景下泛化能力更强
- 适合关注药物设计泛化能力的研究者使用
人工智能在科学领域的进展使蛋白质-配体结合亲和力预测成为可能。然而,现有模型过度依赖简化数据集,无法反映真实生物条件下带有突变、插入、缺失及磷酸化等修饰的蛋白质。本文通过整合4,032个涉及上述修饰的激酶-配体对,构建了完整的、含修饰信息的DAVIS数据集。基于此,提出三种基准测试设置:增强数据预测、野生型到修饰泛化、少样本修饰泛化,用于评估模型在真实生物条件下的鲁棒性。对无对接与有对接方法的广泛评估表明,有对接模型在零样本场景下泛化性能更优;而无对接模型易过拟合野生型蛋白,难以处理未见修饰,但在少量修饰样本微调后表现显著提升。该数据集与基准测试已开源,为开发可泛化至蛋白修饰的模型提供重要基础,助力精准药物研发。
原文摘要 · Abstract (English)
Advancements in AI for science unlocks capabilities for critical drug discovery tasks such as protein-ligand binding affinity prediction. However, current models overfit to existing oversimplified datasets that does not represent naturally occurring and biologically relevant proteins with modifications. In this work, we curate a complete and modification-aware version of the widely used DAVIS dataset by incorporating 4,032 kinase-ligand pairs involving substitutions, insertions, deletions, and phosphorylation events. This enriched dataset enables benchmarking of predictive models under biologically realistic conditions. Based on this new dataset, we propose three benchmark settings-Augmented Dataset Prediction, Wild-Type to Modification Generalization, and Few-Shot Modification Generalization-designed to assess model robustness in the presence of protein modifications. Through extensive evaluation of both docking-free and docking-based methods, we find that docking-based model generalize better in zero-shot settings. In contrast, docking-free models tend to overfit to wild-type proteins and struggle with unseen modifications but show notable improvement when fine-tuned on a small set of modified examples. We anticipate that the curated dataset and benchmarks offer a valuable foundation for developing models that better generalize to protein modifications, ultimately advancing precision medicine in drug discovery. The benchmark is available at: https://github.com/ZhiGroup/DAVIS-complete
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。