arXiv:2605.01681cs.LGq-bio.BM2026-05

对比多种分子对接方法,发现机器学习重排序能显著提升虚拟筛选效果。

Benchmarking Single-Pose Docking, Consensus Rescoring, and Supervised ML on the LIT-PCBA Library: A Critical Evaluation of DiffDock, AutoDock-GPU, GNINA, and DiffDock-NMDN

  • 用经典对接+机器学习重排序,提升筛选精度
  • 监督学习重排序使关键指标提升110%至4.49
  • 适合需要高精度筛选的药物研发人员

虚拟筛选性能高度依赖对接与打分方法的选择。尽管DiffDock和NMDN等基于AI的工具在基准测试中表现优异,但其在真实实验数据上的实用性尚不明确。本文在包含15个靶点、578,295个配体-靶标对的LIT-PCBA库上进行了大规模评估,比较了AutoDock-GPU与DiffDock的构象生成能力,并结合GNINA和NMDN进行重打分。结果表明,AutoDock-GPU搭配GNINA(AutoDock-GNINA)是表现最佳的单一方法,中位数EF1%达2.14。而基于DiffDock的方法在复杂靶点(如OPRK1)上表现较差。精心设计的共识排序提升了鲁棒性,但未超越最优单模型。监督学习重排序带来最大提升,中位数EF1%达4.49,较AutoDock-GNINA提升110%。研究指出,即使最优的古典方法与机器学习结合,也仅在真实基准上提供有限的早期富集。结论是:无单一方法适用于所有靶点,当前最实用的方案是经验证的低成本组合策略,配合监督学习重排序。

原文摘要 · Abstract (English)

Virtual screening performance depends heavily on the chosen docking and scoring methods. Recent AI-based tools such as DiffDock and NMDN have reported strong benchmark results, but their practical utility on realistic, experimentally-derived datasets remains unclear. Here we perform a large-scale evaluation on the LIT-PCBA library (15 targets, 578,295 ligand-target pairs with experimentally confirmed actives and inactives). We compare AutoDock-GPU and DiffDock for pose generation, followed by rescoring with GNINA and NMDN. We further evaluate rank-based consensus strategies and supervised machine learning models trained on docking features. GNINA rescoring of AutoDock-GPU poses (AutoDock-GNINA) emerged as the strongest single method with a median EF1% of 2.14. DiffDock-based approaches underperformed relative to AutoDock-GNINA, particularly on challenging targets such as OPRK1. Carefully designed consensus ranking improved robustness but did not surpass the best single scorer. Supervised ML re-ranking delivered the largest gains, achieving a median EF1% of 4.49 (+110% over AutoDock-GNINA). Our results highlight that even the best classical+ML hybrid workflows provide only modest early enrichment on realistic benchmarks. We conclude that no single docking method dominates across targets and that rigorously validated, cost-effective combinations with supervised re-ranking currently offer the most practical value for virtual screening.

虚拟筛选分子对接机器学习药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。