用神经网络加速超大规模化合物库搜索,1分钟内精准找到最优候选分子。
APEX: Approximate-but-exhaustive search for ultra-large combinatorial synthesis libraries
- 基于神经网络代理模型,利用化合物库结构特性实现快速近似穷举搜索。
- 在百万级化合物库上实现1分钟内完成精确的top-k检索,精度优于现有方法。
- 适合药物研发中需频繁调整筛选条件的场景,支持高效迭代优化。
按需生成的组合合成库(CSLs)如Enamine REAL显著推动了药物发现进程。然而,其规模高达数十亿化合物,给虚拟筛选带来挑战:在有限算力预算下,传统方法通常只能评估不足0.1%的化合物,可能遗漏高分候选。且筛选过程中目标或约束变化时,现有算法难以复用计算资源。本文提出近似但穷尽的搜索协议APEX,通过神经网络代理模型利用化合物库结构特征预测评分与约束,使在消费级GPU上实现全枚举成为可能,可在1分钟内完成精确的近似top-k检索。为验证性能,我们构建了一个包含超过1000万化合物的基准库,所有化合物均标注了在五个临床相关靶点上的对接分数及RDKit测量的理化性质,可针对任意目标和约束条件获取真实最优解,用于对比不同算法表现。实验表明,APEX在检索精度与运行时间上均显著优于其他方法。
原文摘要 · Abstract (English)
Make-on-demand combinatorial synthesis libraries (CSLs) like Enamine REAL have significantly enabled drug discovery efforts. However, their large size presents a challenge for virtual screening, where the goal is to identify the top compounds in a library according to a computational objective (e.g., optimizing docking score) subject to computational constraints under a limited computational budget. For current library sizes -- numbering in the tens of billions of compounds -- and scoring functions of interest, a routine virtual screening campaign may be limited to scoring fewer than 0.1% of the available compounds, leaving potentially many high scoring compounds undiscovered. Furthermore, as constraints (and sometimes objectives) change during the course of a virtual screening campaign, existing virtual screening algorithms typically offer little room for amortization. We propose the approximate-but-exhaustive search protocol for CSLs, or APEX. APEX utilizes a neural network surrogate that exploits the structure of CSLs in the prediction of objectives and constraints to make full enumeration on a consumer GPU possible in under a minute, allowing for exact retrieval of approximate top-k sets. To demonstrate APEX's capabilities, we develop a benchmark CSL comprised of more than 10 million compounds, all of which have been annotated with their docking scores on five medically relevant targets along with physicohemical properties measured with RDKit such that, for any objective and set of constraints, the ground truth top-k compounds can be identified and compared against the retrievals from any virtual screening algorithm. We show APEX's consistently strong performance both in retrieval accuracy and runtime compared to alternative methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。